Skip to content

LLM Pareto Frontier — price vs intelligence, updated monthly

Living report · updated monthly · last updated
Edition: 2026-08 (latest) · all editions: 2026-08 · 2026-07

A model is Pareto-efficient if no other model is both cheaper and smarter. This report plots the current flagship LLMs — the open-weight models we serve and the closed models from Anthropic, OpenAI and Google — on output price versus the Artificial Analysis Intelligence Index, and computes the frontier from the data. Everything below the line is dominated: a strictly better deal exists.

Method: all numbers come from the same source on the same day — Artificial Analysis model pages (Intelligence Index, vendor list prices). Each edition is archived so you can watch the frontier move over time.

$0$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index Claude Opus 5 GPT-5.6 Sol Claude Fable 5 Opus 4.8 GPT-5.5 Sonnet 5 GPT-5.4 Gemini 3.5 Flash Kimi K3 Qwen3.8 Max GLM 5.2 MiniMax M3 Kimi K2.6 Kimi K2.7 DeepSeek V4 Flash GLM 5.3 GLM-5.3-Flash the top open model is now 4 pointsoff the top score — for 40% less Open weights (we serve them) Proprietary Pareto frontier

Reading left to right: GLM-5.3-Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5.

  • Scores across the whole board moved this edition — that’s Artificial Analysis, not the models. AA recalibrated its Intelligence Index to v4.1.1 (nine evaluations, GDPval-AA v2 and τ³-Banking among them) and every tracked model shifted up 1–5 points at unchanged prices. Compare positions, not raw scores, across editions.
  • GLM-5.3-Flash is the move of this refresh (2026-08-31): index 57 at $0.50/M output. An 18B-active, MIT-licensed, natively multimodal MoE that lands straight on the frontier and collapses its entire cheap half into one point: DeepSeek V4 Flash, MiniMax M3 and Kimi K2.7 are all dominated by it, and it matches Claude Opus 4.8 (57 at $25) at 50× less. Full analysis: GLM-5.3-Flash: specs, benchmarks, pricing.
  • GLM 5.3 was the move of the 08-30 refresh: index 60 at $4.40/M output — Kimi K3’s score at 3.4× less. Z.ai open-weighted it on August 28 (GLM-5.3 License) and it replaced GLM 5.2 in our Frontier Pool two days later. On the chart it dominates Qwen3.8 Max (58 at $6), Claude Sonnet 5 (55 at $10) and its own predecessor GLM 5.2 (53 at $4.40), and it matches Kimi K3 (60 at $15) — so K3 drops off the frontier line without losing a point.
  • GPT-5.6 Sol cut its output price $30 → $20 and re-enters the frontier at 61: the last two steps are now GLM 5.3 (60, $4.40) → GPT-5.6 Sol (61, $20) → Claude Opus 5 (63, $25). Two closed models at the top, one point apart, at 4.5–5.7× the open price.
  • The open-vs-closed gap stays at 3 points (GLM 5.3 and Kimi K3 at 60 vs Claude Opus 5’s 63) — per Artificial Analysis, the narrowest since the GLM-5 release in February — but the cheapest way to buy those 60 points is now $4.40/M, not $15.
  • Editorial note (2026-09-10): DeepSeek V4.1 Flash is pending an AA index. DeepSeek released a new generation of its Flash model on September 10 — 552B MoE, MIT weights, natively multimodal, at a lower list price ($0.30/$1.20 peak) — and it replaced V4 Flash in our Core Pool the same day. Artificial Analysis has not published an Intelligence Index score for it, so nothing on this chart moves: the plotted DeepSeek point remains V4 Flash, at its own measured index and price, until an independent score exists. It is on the collection watchlist and enters the tracked set the moment AA scores it.
  • DeepSeek V4 Flash’s new peak list pricing ($0.14/$0.28 → $0.44/$1.32) had put MiniMax M3 back on the frontier as the cheap anchor on 08-14 — a reign that lasted until GLM-5.3-Flash’s arrival swept both off the line.
  • Below $20 per million output tokens, the frontier is now just two models, both open-weight and both Z.ai’s (GLM-5.3-Flash and GLM 5.3). GLM 5.3 is served in our pools; Flash is under evaluation (live status). Kimi K3 and Qwen3.8 Max, one row behind the line, are served too.
  • The near-vertical cliff at the top is back, in a new place: GLM 5.3 (60) → GPT-5.6 Sol (61) is 1 point for 4.5× the price. Claude Fable 5 (62 at $50) remains dominated by Opus 5.
ModelAA IndexInput $/MOutput $/MWeightsPareto-efficient
Claude Opus 563$5.00$25.00closed
Claude Fable 562$10.00$50.00closeddominated by Claude Opus 5
GPT-5.6 Sol61$4.00$20.00closed
GLM 5.360$1.40$4.40open
Kimi K360$3.00$15.00openmatched by GLM 5.3 at 3.4× less
Qwen3.8 Max58$2.00$6.00opendominated by GLM 5.3
GLM-5.3-Flash57$0.15$0.50open
Opus 4.857$5.00$25.00closedmatched by GLM-5.3-Flash at 50.0× less
GPT-5.556$5.00$30.00closeddominated by Claude Opus 5
Sonnet 555$2.00$10.00closeddominated by GLM 5.3
GLM 5.253$1.40$4.40opendominated by GLM 5.3
GPT-5.453$2.50$15.00closedmatched by GLM 5.2 at 3.4× less
DeepSeek V4 Flash52$0.44$1.32opendominated by GLM-5.3-Flash
Gemini 3.5 Flash52$1.50$9.00closedmatched by DeepSeek V4 Flash at 6.8× less
MiniMax M345$0.30$1.20opendominated by GLM-5.3-Flash
Kimi K2.645$0.95$4.00openmatched by MiniMax M3 at 3.3× less
Kimi K2.743$0.95$4.00opendominated by GLM-5.3-Flash

Each edition’s frontier is archived; this chart overlays them, current on top. Over time it shows the defining dynamic of this market: the frontier sliding down (cheaper) and right-side-up (smarter) — driven almost entirely by open-weight releases.

$0.00$10$20$30$40$50 Output price, $ per 1M tokens 4045505560 AA Intelligence Index 2026-07 2026-08 (current) Current frontier Previous editions (older = fainter) Open Closed

The Intelligence Index is one composite — task-specific rankings differ. Kimi K2.6 sits off-frontier here yet holds the best open SWE-bench Verified score (80.2); for agentic coding the ranking flips — see the Which-LLM guide for tier-fair matchups. Prices are vendor list prices for the reasoning variants Artificial Analysis evaluates; open-weight prices vary by host. MiMo V2.5 and DeepSeek V4.1 Flash, which we also serve, have no index entry yet and are excluded rather than estimated — the plotted DeepSeek point is the V4 Flash generation it replaced. Index points aren’t linear in value.

  • 2026-09-10 (no frontier change)DeepSeek V4.1 Flash ships and replaces V4 Flash in our Core Pool: 552B MoE (8B/16B active), causal encoder–decoder, native vision, MIT weights, $0.30/$1.20 peak list — cheaper than the model it replaces. Artificial Analysis has not scored it, so it is added to the watchlist and no point on this chart moves; the DeepSeek point stays V4 Flash until an independent index exists.

  • 2026-08 refresh (2026-08-31)GLM-5.3-Flash joins the tracked set (57, $0.15/$0.50) and lands straight on the frontier: Z.ai’s 18B-active, MIT-licensed, natively multimodal MoE matches Claude Opus 4.8 at 50× less and dominates DeepSeek V4 Flash, MiniMax M3 and Kimi K2.7 — the frontier’s whole cheap half collapses into one point. Frontier: GLM-5.3-Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. Analysis: GLM-5.3-Flash: specs, benchmarks & pricing.

  • 2026-08 refresh (2026-08-30)GLM 5.3 joins the tracked set (60, $1.40/$4.40) two days after its open weights shipped under the GLM-5.3 License, and enters the frontier at once: it dominates Qwen3.8 Max, Sonnet 5 and GLM 5.2, and matches Kimi K3 at 3.4× less — K3 leaves the line. GPT-5.6 Sol’s output price drops $30 → $20 and it re-enters the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.3 → GPT-5.6 Sol → Claude Opus 5. GLM 5.2 stays tracked for history; it was upgraded in place to 5.3 in our Frontier Pool.

  • 2026-08 refresh (2026-08-14) — three moves in one collection. AA recalibrated its index to v4.1.1: every tracked score shifted up 1–5 points at unchanged prices — cross-edition score jumps around this date are methodology, not model changes. Qwen3.8 Max joins the tracked set (58, $2/$6) now that its weights shipped — it enters the frontier and pushes Sonnet 5 off it, leaving Opus 5 as the only closed model on the line. And DeepSeek V4 Flash’s peak list pricing took effect ($0.14/$0.28 → $0.44/$1.32), bringing MiniMax M3 back onto the frontier. Frontier: MiniMax M3 → DeepSeek V4 Flash → GLM 5.2 → Qwen3.8 Max → Kimi K3 → Claude Opus 5.

  • 2026-08 (2026-08-01) — release-week update: the V4-Flash-0731 retrain lifts DeepSeek V4 Flash 40 → 50 at unchanged prices ($0.14/$0.28) — a 10-point single-model jump. It now matches Gemini 3.5 Flash at 32× lower output price and pushes MiniMax M3 off the frontier (dominated). Frontier: DeepSeek V4 Flash → GLM 5.2 → Sonnet 5 → Kimi K3 → Claude Opus 5.

  • 2026-07 refresh (2026-07-28) — release-week update: Kimi K3 (57, $3/$15), Claude Opus 5 (61, $5/$25) and GPT-5.6 Sol (59, $5/$30) added to the tracked set. K3 and Opus 5 enter the frontier; Fable 5 and Opus 4.8 leave it (both dominated by Opus 5). Sonnet 5’s output price dropped $15 → $10. The open-vs-closed gap is now 4 index points — the narrowest since GLM-5 (February).

  • 2026-07 (first edition) — baseline. Frontier: DeepSeek V4 Flash, MiniMax M3, GLM 5.2, Sonnet 5, Opus 4.8, Fable 5. Notable context at launch: GLM 5.2 (released mid-June) is the top-scoring open-weights model in index history; AA’s v4.1 recalibration lowered scores across the board vs the April v4.0 figures.


GLM 5.3 — the frontier’s top open model — plus Kimi K3 and Qwen3.8 Max just behind the line are served flat-rate in our pools — Flagship from $169.15/mo, Frontier from $60.35/mo, Core from $15.29/mo — where the marginal token costs zero during your reserved hours. GLM-5.3-Flash is under evaluation.