The Price of Intelligence: Open-Weight vs Paid Frontier Models in 2026
The usual way people compare open-source and paid AI models is to line up benchmark scores and declare a winner. That comparison is almost useless, because it answers a question nobody is actually asking. You don't choose a model in a vacuum — you choose it against a price. The real question is not "which model is smartest" but "how much intelligence does a dollar buy, and where does paying more stop being worth it."
So this post has one central chart, and everything else supports it. I took every frontier model I could get consistent numbers for — the paid flagships (Claude Opus 4.8, Claude Fable 5, OpenAI's GPT-5.6 family, xAI's Grok 4.5) and the top open-weight models (DeepSeek V4 Pro, GLM-5.2, Kimi K2.6, Qwen3.5) — and plotted what each one charges against how capable it is.
The answer is more interesting than either camp's slogan. Open weights have not "caught up," and they have not "stayed behind." They have caught up in one specific place — the middle — and the paid premium has retreated to the top.
The one chart
Chart note (22 July): positions for models without figures sourced in this post are approximate — read these charts as the shape of the market, not exact coordinates.
Read it like an economist reads a cost curve. The vertical axis is the Artificial Analysis Intelligence Index — a composite that, in its current version, weights agentic and coding tasks heavily. The horizontal axis is output-token price, on a log scale so that $4 and $50 both read clearly. The dashed line is the efficient frontier: the best intelligence available at each price. Anything below and to the right of it is paying more for less.
Three things fall out immediately.
The open models own the cheap end of the frontier. DeepSeek V4 Pro and GLM-5.2 sit right on the efficient line — you cannot get their level of capability for less money from anyone, open or closed. GLM-5.2 scores 51.1 on the index at roughly $4.40 per million output tokens.
An open model quietly dominates a paid one. GLM-5.2's 51.1 essentially matches GPT-5.6 Luna's 51.2 — at a lower price. That is not "open is competitive," that is an open-weight model beating a frontier lab's product on both axes at once. If you were reaching for Luna for a cost-sensitive job, GLM-5.2 is strictly the better deal.
But the ceiling is closed-only. Look at the top of the chart. Everything above index 56 — GPT-5.6 Sol at 59, Claude Fable 5 at 60 — is paid, and there is no open-weight model anywhere near it. The best open model tops out around 51. That gap is real, and it is where your money still goes.
Update (18 July): Kimi K3's release moved the open-weights ceiling to index 57 — see the K3 post for the new picture.
Open weights didn't catch the frontier. They caught the mid-tier — and made the mid-tier cheap. What you now pay a premium for is the last nine index points at the very top.
The middle is a tie; the top is not
Stack the same intelligence scores vertically and the structure of the market is obvious.
There is a band from roughly index 51 to 55 — GPT-5.6 Luna, GLM-5.2, Grok 4.5, GPT-5.6 Terra — where open and paid models are genuinely interchangeable on capability. Above that band, from 56 to 60, it is paid-only: Claude Opus 4.8, GPT-5.6 Sol, and Claude Fable 5, with nothing open to challenge them.
One important caveat, stated plainly because it cuts against the open-source story: this index is weighted toward agentic workloads — long chains of tool calls, multi-step autonomy. That is exactly the kind of task where the paid labs have poured in the most reinforcement learning, and exactly where open weights lag most. On narrower measures the gap shrinks dramatically: DeepSeek V4 Pro posts around 80.6% on SWE-bench Verified, within a couple of points of the paid frontier on raw coding. So "the gap is nine points" is true for autonomous agents and misleading for everything else. Which benchmark you weight is which decision you're making.
What a dollar actually buys
Flip the axes and ask the buyer's question directly: per dollar of output tokens, how many index points do you get?
DeepSeek V4 Pro returns about 12.7 index points per dollar; Claude Fable 5 returns 1.2. That is a tenfold difference in raw efficiency. Every open model beats every paid model on this measure, and it is not close.
This is the number that should reframe how you budget. If your workload lives in that interchangeable 51–55 band — most extraction, classification, summarization, first-draft generation, and a lot of everyday coding — you are paying a 5-to-10x premium for index points you would have gotten anyway. The paid tax only earns its keep when you genuinely need the top of the chart.
One honest asterisk on this chart: it measures sticker efficiency. It assumes every model spends the same number of tokens to finish a task, and they don't — which brings us to the one place the paid side fights back.
The catch: sticker price isn't task price
Cheaper per token does not always mean cheaper per job, because verbose models burn more tokens to reach the same answer. The sharpest example runs the other way — a paid model that is cheap and terse.
Grok 4.5 is priced at $2 input / $6 output, and xAI reports it solves a task with roughly 4.2x fewer tokens than Opus 4.8. Multiply the two effects together and a finished task can cost a small fraction of the Opus equivalent. Token efficiency is a hidden axis that sticker-price charts miss entirely — and it can flip the ranking. The same caution applies to open models: some of them are notably verbose, which quietly erodes their per-token advantage on real jobs. Always measure cost per completed task on your own workload, not cost per token on a price sheet.
The map, and the decision
Put price and intelligence together as a positioning map and the strategy writes itself.
Open weights own the value zone. The paid flagships own the premium tier. And a couple of paid mid-tier models (Luna, Terra) sit awkwardly in between, now undercut on capability-per-dollar by open alternatives. So:
- High-volume, cost-sensitive, everyday work — extraction, classification, routine coding, first drafts. Use an open model. DeepSeek V4 Pro or GLM-5.2 give you paid-mid-tier capability at a fraction of the price, and you can self-host for data control. This is where most tokens are actually spent, and paying frontier prices here is waste.
- Hard, autonomous, long-horizon agentic work — the multi-step tool-use tasks the index rewards. Pay for the ceiling: GPT-5.6 Sol or Claude Fable 5. The nine-point gap is real here, and on work where a failure is expensive, it is worth every cent.
- Cost-sensitive but latency- or efficiency-critical — look hard at Grok 4.5. Its token efficiency can make it the cheapest per task even against models with lower sticker prices.
What you're actually paying for
When you pay the frontier premium in mid-2026, you are not buying raw intelligence in the abstract — open weights have made that cheap. You are buying four specific things: the last nine index points at the very top of the capability curve, the agentic reliability that comes with them, the managed API and tooling around them, and someone else's responsibility for safety and uptime.
For a large fraction of real workloads, none of those four is worth a 5-to-10x markup, and an open model is simply the correct engineering choice. For the hardest autonomous work, all four still matter, and the paid frontier earns its price. The mistake — the expensive, common mistake — is paying premium-tier prices for value-zone work out of habit, because the benchmark headline said the paid model was "better."
Better at what, and worth how much more? Those are the only two questions, and the chart at the top answers both.
An interactive version
The price-versus-intelligence map works best when you can poke at it — filter open versus paid, re-rank by value per dollar, and read each model's exact coordinates on hover. A static image can't do that, so I built a live version: open The Price of Intelligence — the interactive frontier map.
The map ranks models by a third-party intelligence index against what a task actually costs. For how these models behave on identical work — the same coding, reading and drawing tasks, with every prompt and each model's raw output shown — see the model comparison page, which is measured first-hand and grows as new models are tested.
Sources & caveats
Every number here traces to one of the sources below. Two honest caveats. First, third-party benchmark aggregators disagree at the margins and open-model version numbers move fast (GLM-5.1 vs 5.2, Kimi K2.6 vs K2.7) — I used a single consistent leaderboard snapshot for the intelligence scores so the comparison is internally coherent, but treat individual figures as accurate-as-of-July-2026, not eternal. Second, the intelligence index is agentic-weighted; a coding-only or reasoning-only index would narrow the open-vs-paid gap considerably, and I've said so where it matters.
- Artificial Analysis — Intelligence Index (the capability axis, methodology)
- BenchLM — Artificial Analysis Index leaderboard, July 2026 (per-model index scores)
- BenchLM — Best open-source LLM 2026 (open-model rankings, SWE-bench)
- DeepInfra — DeepSeek V4 Pro pricing guide and xAI — Grok 4.5 announcement (Grok pricing + token-efficiency, vendor page)
- Artificial Analysis — GPT-5.6 has landed (GPT-5.6 tier scores and pricing)
- xAI — Grok 4.5 announcement (token-efficiency chart, vendor page)