Model Face-Off
← back to AI Signal

How steady are these results?

A single attempt can make a coin-flip behavior look like a fixed trait, in either direction — that's what the four runs per task are for. The table below is the complete record: all twenty-five tasks, every model, every run. Each cell is how many of that model’s runs were a full pass (every hidden check) on that task; the small line in the Task 19 row breaks that task’s runs down score by score. Task 24, the free drawing, has no hidden checks; its row shows elements found on my checklist (out of 12) and my rank.

A companion check: two tasks (the long-document recall and the closures/caches task) were also re-run for the three Claude models in a fully sealed setup — the model got the prompt text and nothing else, with no tools, no files, and no possible route to the checks or answer keys. All three scored exactly the same sealed as unsealed — evidence, not proof, that the numbers here reflect real capability rather than lucky access.

TaskSonnet 5Fable 5Qwen3.83.6 FlashOpus 4.8Grok 4.5DeepSeek3.5 Flash-LiteOpus 5V4 FlashSolTerraLuna
Task 1coding4/44/44/44/44/44/44/44/44/44/44/44/42/4
Task 2coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 3coding4/43/44/40/41/43/42/40/43/41/44/40/42/4
Task 4coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 5coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 6coding4/44/44/44/44/44/44/42/44/44/44/44/44/4
Task 7coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 8coding4/41/43/34/44/42/40/40/44/41/44/41/42/4
Task 9coding4/44/44/44/44/44/44/44/44/44/44/42/44/4
Task 10coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 11coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 12coding1/44/41/42/43/42/42/40/44/44/44/44/44/4
Task 13coding4/44/44/44/44/44/44/44/44/44/44/44/43/4
Task 14coding4/44/44/44/44/44/44/41/44/44/44/44/44/4
Task 15coding4/44/44/44/44/44/44/42/44/44/44/44/43/4
Task 16coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 17coding4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 18coding4/44/44/44/44/44/44/44/44/44/42/42/44/4
Task 19language0/40/40/40/40/40/40/43/40/40/40/40/40/4
Task 20language4/44/43/44/44/44/44/41/44/44/44/41/43/4
Task 21language4/44/44/44/44/44/44/40/44/43/44/42/44/4
Task 22reading4/44/44/44/44/44/44/44/44/44/44/44/44/4
Task 23reading4/44/44/44/44/44/44/40/44/44/44/43/44/4
Task 24drawing7/12my rank: 5th9/12my rank: 2nd12/12my rank: 1st11/12my rank: 1st8/12my rank: 4th5/12my rank: 6th9/12my rank: 3rd10/12my rank: 4th12/12my rank: 1stcollectedrank pendingcollectedrank pendingcollectedrank pendingcollectedrank pending
Task 25drawing4/44/44/43/44/44/44/44/44/44/44/44/44/4

Reading the table: a cell is the count of runs that passed every hidden check for that task. — means no runs have been done yet for that model on that task. Every cell rests on four complete runs, with one stated exception: Qwen's Task 8 is three. Three samples had earlier been lost to a saving problem on our side and were left out of the totals; they were re-collected on 2026-07-29, so nothing is excluded any more — and none of the published totals moved when they came back.

Behind the partial cells: Flash-Lite’s contradiction-hunt runs (Task 23) always find the device and timeout conflicts but dismiss at least one of the other two; its acrostic misses (Task 20) usually drop the seven-words-per-line rule, while Qwen’s one acrostic miss writes two six-word lines; Gemini 3.6 Flash’s one strict-drawing miss (Task 25) is a 9 of 10.