How steady are these results?
A single attempt can make a coin-flip behavior look like a fixed trait, in either direction — that's what the four runs per task are for. The table below is the complete record: all twenty-five tasks, every model, every run. Each cell is how many of that model’s runs were a full pass (every hidden check) on that task; the small line in the Task 19 row breaks that task’s runs down score by score. Task 24, the free drawing, has no hidden checks; its row shows elements found on my checklist (out of 12) and my rank.
A companion check: two tasks (the long-document recall and the closures/caches task) were also re-run for the three Claude models in a fully sealed setup — the model got the prompt text and nothing else, with no tools, no files, and no possible route to the checks or answer keys. All three scored exactly the same sealed as unsealed — evidence, not proof, that the numbers here reflect real capability rather than lucky access.
| Task | Sonnet 5 | Fable 5 | Qwen3.8 | 3.6 Flash | Opus 4.8 | Grok 4.5 | DeepSeek | 3.5 Flash-Lite | Opus 5 | V4 Flash | Sol | Terra | Luna |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Task 1coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 2/4 |
| Task 2coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 3coding | 4/4 | 3/4 | 4/4 | 0/4 | 1/4 | 3/4 | 2/4 | 0/4 | 3/4 | 1/4 | 4/4 | 0/4 | 2/4 |
| Task 4coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 5coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 6coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 2/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 7coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 8coding | 4/4 | 1/4 | 3/3 | 4/4 | 4/4 | 2/4 | 0/4 | 0/4 | 4/4 | 1/4 | 4/4 | 1/4 | 2/4 |
| Task 9coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 2/4 | 4/4 |
| Task 10coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 11coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 12coding | 1/4 | 4/4 | 1/4 | 2/4 | 3/4 | 2/4 | 2/4 | 0/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 13coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 3/4 |
| Task 14coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 1/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 15coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 2/4 | 4/4 | 4/4 | 4/4 | 4/4 | 3/4 |
| Task 16coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 17coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 18coding | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 2/4 | 2/4 | 4/4 |
| Task 19language | 0/4 | 0/4 | 0/4 | 0/4 | 0/4 | 0/4 | 0/4 | 3/4 | 0/4 | 0/4 | 0/4 | 0/4 | 0/4 |
| Task 20language | 4/4 | 4/4 | 3/4 | 4/4 | 4/4 | 4/4 | 4/4 | 1/4 | 4/4 | 4/4 | 4/4 | 1/4 | 3/4 |
| Task 21language | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 0/4 | 4/4 | 3/4 | 4/4 | 2/4 | 4/4 |
| Task 22reading | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
| Task 23reading | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 0/4 | 4/4 | 4/4 | 4/4 | 3/4 | 4/4 |
| Task 24drawing | 7/12my rank: 5th | 9/12my rank: 2nd | 12/12my rank: 1st | 11/12my rank: 1st | 8/12my rank: 4th | 5/12my rank: 6th | 9/12my rank: 3rd | 10/12my rank: 4th | 12/12my rank: 1st | collectedrank pending | collectedrank pending | collectedrank pending | collectedrank pending |
| Task 25drawing | 4/4 | 4/4 | 4/4 | 3/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 | 4/4 |
Reading the table: a cell is the count of runs that passed every hidden check for that task. — means no runs have been done yet for that model on that task. Every cell rests on four complete runs, with one stated exception: Qwen's Task 8 is three. Three samples had earlier been lost to a saving problem on our side and were left out of the totals; they were re-collected on 2026-07-29, so nothing is excluded any more — and none of the published totals moved when they came back.
Behind the partial cells: Flash-Lite’s contradiction-hunt runs (Task 23) always find the device and timeout conflicts but dismiss at least one of the other two; its acrostic misses (Task 20) usually drop the seven-words-per-line rule, while Qwen’s one acrostic miss writes two six-word lines; Gemini 3.6 Flash’s one strict-drawing miss (Task 25) is a 9 of 10.