The results
| Model | Coding Tasks 1–18 · average of 4 runs · worst |
Fact recall Task 22 |
Spotting contradictions Task 23 |
Free drawing Task 24 · elements of 12 · my rank |
Strict drawing Task 25 · rules of 10 · elements of 7 · my rank |
Reasoning & language Tasks 19–21 · average · worst |
|---|---|---|---|---|---|---|
| Claude Opus 5Anthropic · default · via CLI | 123.2 / 124 · worst 121 | 5 / 5 | 4 / 4 | 12/12 · 1st | 10 / 10 · 7/7 · 1st | 12.0 / 13 · worst 12 |
| Claude Fable 5Anthropic · High | 122.2 / 124 · worst 119 | 5 / 5 | 4 / 4 | 9/12 · 2nd | 10 / 10 · 7/7 · 2nd | 12.0 / 13 · worst 12 |
| GPT-5.6 SolOpenAI · default · via CLI | 121.5 / 124 · worst 119 | 8 / 8 | 4 / 4 | collected · rank pending | 10 / 10 · rank pending | 12.0 / 13 · worst 12 |
| Claude Opus 4.8Anthropic · High | 120.8 / 124 · worst 115 | 5 / 5 | 4 / 4 | 8/12 · 4th | 10 / 10 · 7/7 · 4th | 11.0 / 13 · worst 11 |
| DeepSeek V4-FlashOpenCode · free route · pure clean-context CLI | 120.8 / 124 · worst 119 | 8 / 8 | 4 / 4 | collected · rank pending | 10 / 10 · rank pending | 11.5 / 13 · worst 10 |
| Claude Sonnet 5Anthropic · Medium | 120.2 / 124 · worst 119 | 5 / 5 | 4 / 4 | 7/12 · 5th | 10 / 10 · 7/7 · 4th | 11.8 / 13 · worst 11 |
| Qwen3.8-Max-PreviewAlibaba · Think | 120.2 / 124 · worst 119 | 5 / 5 | 4 / 4 | 12/12 · 1st | 10 / 10 · 7/7 · 1st | 11.8 / 13 · worst 11 |
| Grok 4.5xAI · Fast | 119.5 / 124 · worst 112 | 5 / 5 | 4 / 4 | 5/12 · 6th | 10 / 10 · 5/7 · 5th | 12.0 / 13 · worst 12 |
| GPT-5.6 LunaOpenAI · default · via CLI | 119.2 / 124 · worst 112 | 8 / 8 | 4 / 4 | collected · rank pending | 10 / 10 · rank pending | 11.8 / 13 · worst 11 |
| Gemini 3.6 FlashGoogle · default · via API | 118.5 / 124 · worst 116 | 5 / 5 | 4 / 4 | 11/12 · 1st | 9 / 10 · 7/7 · 3rd | 11.5 / 13 · worst 11 |
| DeepSeekDeepSeek · Expert+DeepThink | 117.2 / 124 · worst 106 | 5 / 5 | 4 / 4 | 9/12 · 3rd | 10 / 10 · 7/7 · 4th | 12.0 / 13 · worst 12 |
| GPT-5.6 TerraOpenAI · default · via CLI | 114.2 / 124 · worst 109 | 8 / 8 | 3 / 4 | collected · rank pending | 10 / 10 · rank pending | 10.5 / 13 · worst 9 |
| Gemini 3.5 Flash-LiteGoogle · default · via API | 109.8 / 124 · worst 104 | 5 / 5 | 2 / 4 | 10/12 · 4th | 10 / 10 · 6/7 · 4th | 10.8 / 13 · worst 9 |
Gemini 3.5 Flash-Lite's fact-recall score comes from an automated keyword scan rather than a human read.
Both drawing tasks carry two grades: a machine count of the elements or rules a picture actually satisfies, and my own rank for how it looks. The strict drawing shows why both are needed — twelve of the thirteen models satisfy all ten rules in every run, and I still rank them from first to last. Identical scores, visibly different pictures. A drawing can satisfy every requirement and still be the worse one.
Reading the table. The headline number is a model's average across four separate runs, with its worst run beside it. Averaging is what makes columns comparable — it does not drift as you run a model more times, and it still works where a column has three runs rather than four. The worst run is the number to read if you plan to take an answer unchecked. Claude Opus 5 leads on both, averaging 123.2 of 124 and never dropping below 121; DeepSeek and Gemini 3.5 Flash-Lite give up the most ground — almost always by repeating one small mistake across runs rather than by failing broadly. Counting whole tasks a model kept clean in every run, out of 18: Opus 5 17, Sonnet 17, Qwen 17, GPT-5.6 Sol 17, Opus 4.8 16, Fable 16, Gemini 3.6 Flash 16, DeepSeek V4-Flash 16, Grok 15, DeepSeek V4 15, GPT-5.6 Terra 14, GPT-5.6 Luna 13, Gemini 3.5 Flash-Lite 12. But the newest and strongest column still fails the sentence-word-count question in every single run: twelve of the thirteen models answer it wrong all four times, and the one model that ever gets it right is the weakest overall. Every model finds all four planted contradictions with zero false alarms except Gemini 3.5 Flash-Lite, which never finds more than three, and one GPT-5.6 Terra run that found three of four. Fifteen of the twenty-four machine-graded tasks carry at least one of these gaps — the run-by-run record shows them run by run. None of it is about which model reasons hardest; it's about where each one's carefulness runs out, and how consistently.
How to read the ordering
The table is ordered by the average run, not the worst. That is a deliberate change. A worst-run score is the minimum of however many times I ran the model, so it keeps sliding downward the more carefully a model is measured — scoring this same data as three runs instead of four lifts DeepSeek by nearly three points. A number that moves with the sample size is not a property of the model, and it cannot be compared against a column run a different number of times.
The worst run is still worth publishing, and it is still in every cell, because it answers a different question: not how good is this model but how bad does it get. Set the two side by side and the distance between them is the interesting part:
| Model | Worst run | Average run | Gap |
|---|---|---|---|
| Claude Opus 5 | 121 / 124 | 123.2 / 124 | −2.2 |
| Claude Fable 5 | 119 / 124 | 122.2 / 124 | −3.2 |
| GPT-5.6 Sol | 119 / 124 | 121.5 / 124 | −2.5 |
| Claude Opus 4.8 | 115 / 124 | 120.8 / 124 | −5.8 |
| DeepSeek V4-Flash | 119 / 124 | 120.8 / 124 | −1.8 |
| Claude Sonnet 5 | 119 / 124 | 120.2 / 124 | −1.2 |
| Qwen3.8-Max-Preview | 119 / 124 | 120.2 / 124 | −1.2 |
| Grok 4.5 | 112 / 124 | 119.5 / 124 | −7.5 |
| GPT-5.6 Luna | 112 / 124 | 119.2 / 124 | −7.2 |
| Gemini 3.6 Flash | 116 / 124 | 118.5 / 124 | −2.5 |
| DeepSeek | 106 / 124 | 117.2 / 124 | −11.2 |
| GPT-5.6 Terra | 109 / 124 | 114.2 / 124 | −5.2 |
| Gemini 3.5 Flash-Lite | 104 / 124 | 109.8 / 124 | −5.8 |
The gap is the column worth staring at: it is the distance between a model's typical work and its bad day. DeepSeek loses eleven points to it and Grok 4.5 seven and a half — both do markedly better work than their worst run implies, and are markedly less dependable. Sonnet 5 and Qwen lose barely one: what they manage once, they manage every time. Claude Opus 4.8 is the clearest case — third on average, sixth on its worst run. The same model, six places apart depending on which question you ask. Ordered by worst run, as this table was until 2026-07-29, it read as the weaker model. It is not weaker; it is streakier.
Two more things before quoting any of this. Four runs is enough to catch a coin-flip but not enough to separate models a few points apart — treat the leaders as a group rather than a ranking; the gap between first and fourth here is roughly one task. And every model ran at its own out-of-the-box setting, which is not the same setting for everyone: Sonnet 5 did this at Medium effort while Fable 5 and Opus 4.8 ran at High.