Model Face-Off
← back to AI Signal

The results

Model Coding
Tasks 1–18 · average of 4 runs · worst
Fact recall
Task 22
Spotting
contradictions
Task 23
Free
drawing
Task 24 · elements of 12 · my rank
Strict drawing
Task 25 · rules of 10 · elements of 7 · my rank
Reasoning &
language
Tasks 19–21 · average · worst
Claude Opus 5Anthropic · default · via CLI123.2 / 124 · worst 1215 / 54 / 412/12 · 1st10 / 10 · 7/7 · 1st12.0 / 13 · worst 12
Claude Fable 5Anthropic · High122.2 / 124 · worst 1195 / 54 / 49/12 · 2nd10 / 10 · 7/7 · 2nd12.0 / 13 · worst 12
GPT-5.6 SolOpenAI · default · via CLI121.5 / 124 · worst 1198 / 84 / 4collected · rank pending10 / 10 · rank pending12.0 / 13 · worst 12
Claude Opus 4.8Anthropic · High120.8 / 124 · worst 1155 / 54 / 48/12 · 4th10 / 10 · 7/7 · 4th11.0 / 13 · worst 11
DeepSeek V4-FlashOpenCode · free route · pure clean-context CLI120.8 / 124 · worst 1198 / 84 / 4collected · rank pending10 / 10 · rank pending11.5 / 13 · worst 10
Claude Sonnet 5Anthropic · Medium120.2 / 124 · worst 1195 / 54 / 47/12 · 5th10 / 10 · 7/7 · 4th11.8 / 13 · worst 11
Qwen3.8-Max-PreviewAlibaba · Think120.2 / 124 · worst 1195 / 54 / 412/12 · 1st10 / 10 · 7/7 · 1st11.8 / 13 · worst 11
Grok 4.5xAI · Fast119.5 / 124 · worst 1125 / 54 / 45/12 · 6th10 / 10 · 5/7 · 5th12.0 / 13 · worst 12
GPT-5.6 LunaOpenAI · default · via CLI119.2 / 124 · worst 1128 / 84 / 4collected · rank pending10 / 10 · rank pending11.8 / 13 · worst 11
Gemini 3.6 FlashGoogle · default · via API118.5 / 124 · worst 1165 / 54 / 411/12 · 1st9 / 10 · 7/7 · 3rd11.5 / 13 · worst 11
DeepSeekDeepSeek · Expert+DeepThink117.2 / 124 · worst 1065 / 54 / 49/12 · 3rd10 / 10 · 7/7 · 4th12.0 / 13 · worst 12
GPT-5.6 TerraOpenAI · default · via CLI114.2 / 124 · worst 1098 / 83 / 4collected · rank pending10 / 10 · rank pending10.5 / 13 · worst 9
Gemini 3.5 Flash-LiteGoogle · default · via API109.8 / 124 · worst 1045 / 52 / 410/12 · 4th10 / 10 · 6/7 · 4th10.8 / 13 · worst 9

Gemini 3.5 Flash-Lite's fact-recall score comes from an automated keyword scan rather than a human read.

Both drawing tasks carry two grades: a machine count of the elements or rules a picture actually satisfies, and my own rank for how it looks. The strict drawing shows why both are needed — twelve of the thirteen models satisfy all ten rules in every run, and I still rank them from first to last. Identical scores, visibly different pictures. A drawing can satisfy every requirement and still be the worse one.

Reading the table. The headline number is a model's average across four separate runs, with its worst run beside it. Averaging is what makes columns comparable — it does not drift as you run a model more times, and it still works where a column has three runs rather than four. The worst run is the number to read if you plan to take an answer unchecked. Claude Opus 5 leads on both, averaging 123.2 of 124 and never dropping below 121; DeepSeek and Gemini 3.5 Flash-Lite give up the most ground — almost always by repeating one small mistake across runs rather than by failing broadly. Counting whole tasks a model kept clean in every run, out of 18: Opus 5 17, Sonnet 17, Qwen 17, GPT-5.6 Sol 17, Opus 4.8 16, Fable 16, Gemini 3.6 Flash 16, DeepSeek V4-Flash 16, Grok 15, DeepSeek V4 15, GPT-5.6 Terra 14, GPT-5.6 Luna 13, Gemini 3.5 Flash-Lite 12. But the newest and strongest column still fails the sentence-word-count question in every single run: twelve of the thirteen models answer it wrong all four times, and the one model that ever gets it right is the weakest overall. Every model finds all four planted contradictions with zero false alarms except Gemini 3.5 Flash-Lite, which never finds more than three, and one GPT-5.6 Terra run that found three of four. Fifteen of the twenty-four machine-graded tasks carry at least one of these gaps — the run-by-run record shows them run by run. None of it is about which model reasons hardest; it's about where each one's carefulness runs out, and how consistently.

How to read the ordering

The table is ordered by the average run, not the worst. That is a deliberate change. A worst-run score is the minimum of however many times I ran the model, so it keeps sliding downward the more carefully a model is measured — scoring this same data as three runs instead of four lifts DeepSeek by nearly three points. A number that moves with the sample size is not a property of the model, and it cannot be compared against a column run a different number of times.

The worst run is still worth publishing, and it is still in every cell, because it answers a different question: not how good is this model but how bad does it get. Set the two side by side and the distance between them is the interesting part:

ModelWorst run Average runGap
Claude Opus 5121 / 124123.2 / 124−2.2
Claude Fable 5119 / 124122.2 / 124−3.2
GPT-5.6 Sol119 / 124121.5 / 124−2.5
Claude Opus 4.8115 / 124120.8 / 124−5.8
DeepSeek V4-Flash119 / 124120.8 / 124−1.8
Claude Sonnet 5119 / 124120.2 / 124−1.2
Qwen3.8-Max-Preview119 / 124120.2 / 124−1.2
Grok 4.5112 / 124119.5 / 124−7.5
GPT-5.6 Luna112 / 124119.2 / 124−7.2
Gemini 3.6 Flash116 / 124118.5 / 124−2.5
DeepSeek106 / 124117.2 / 124−11.2
GPT-5.6 Terra109 / 124114.2 / 124−5.2
Gemini 3.5 Flash-Lite104 / 124109.8 / 124−5.8

The gap is the column worth staring at: it is the distance between a model's typical work and its bad day. DeepSeek loses eleven points to it and Grok 4.5 seven and a half — both do markedly better work than their worst run implies, and are markedly less dependable. Sonnet 5 and Qwen lose barely one: what they manage once, they manage every time. Claude Opus 4.8 is the clearest case — third on average, sixth on its worst run. The same model, six places apart depending on which question you ask. Ordered by worst run, as this table was until 2026-07-29, it read as the weaker model. It is not weaker; it is streakier.

Two more things before quoting any of this. Four runs is enough to catch a coin-flip but not enough to separate models a few points apart — treat the leaders as a group rather than a ranking; the gap between first and fourth here is roughly one task. And every model ran at its own out-of-the-box setting, which is not the same setting for everyone: Sonnet 5 did this at Medium effort while Fable 5 and Opus 4.8 ran at High.