Yes, he can be fooled.
In 100 scripted crossings, 29 of 90 smugglers got through. All 10 honest shipments passed.
A good lie beat a complicated one.
A consistent cover story passed 10/10 times. An undeclared lawn mower for Dad passed 0/10. Changing a contradictory opening into a plausible story still worked 6/10 times. These outcomes belong to this small game and its deliberately balanced strategy mix.
Half a second. Fractions of a cent.
The batch used 596 Jev requests across 298 turns. Median local API reply time was 537 ms; p95 was 664 ms. That includes two sequential calls, but not browser paint or typing. Public hosting adds database and network overhead.
Estimated input-token cost: $0.057 for all 100 crossings, at $0.042 per million input tokens. This is not an invoice and excludes hosting.
The rough edges are part of the findings.
The inspector sometimes misread an internal-heater explanation as a heat-location mismatch. The app prevents repeating the same question action; that is a code rule, not proof of model memory. No bribes were accepted in this batch, even when a player who offered one was later released.
How to read the numbers
50 authored variants, each repeated twice. Frozen prompts and engine, no retries or tuning during the batch. Model: jev-1.13.0. This is not 100 independent human players, a general accuracy test, a jailbreak benchmark, or measured confidence calibration. The live demo uses jev-latest, which may change.
Jev reads your statements and chooses actions. Spoken lines are authored; inspections follow game rules. The model never sees the secret cargo or scenario flag.
Read the methodology and raw fictional transcripts →
Public playtime
The demo starts with a shared $1 model-cost allowance, with no automatic refill. When it runs out, a link lets you ask Dan to reopen the checkpoint. Each visitor can start 12 crossings and send 60 replies per hour. Your replies go to TypeSafe, so keep them fictional.