Skip to content

GLM-5.3: specs, benchmarks, pricing & API options — now served unlimited

Update, August 30, 2026: the review is over — GLM-5.3 is live in our Frontier Pool as glm-5.3. Z.ai published the open weights on August 28 (zai-org/GLM-5.3, under the custom GLM-5.3 License — commercial use allowed; a security-review clause applies only to Model-as-a-Service operators above US$10B revenue), the licensing gate we describe below lifted, and the pool’s GLM slot was upgraded in place: every Frontier subscriber gets GLM-5.3 at the same flat price, and glm-5.2 requests keep working until September 30, 2026. Setup on the model page. The analysis below is as written on launch day.

GLM-5.3 (API id glm-5.3) is Z.ai’s (Zhipu AI) new coding and agentic model, released today, August 14, 2026, under the tagline “Built to Code. Ready for Cyber Defense.” The architecture story is unusual and worth being precise about: GLM-5.3 keeps the same 743B-parameter base model as GLM 5.2 — every reported gain comes from scaled-up post-training alone. On Z.ai’s own benchmark suite that post-training buys a lot: it calls GLM-5.3 the strongest open-weights coding model it has measured, and reports a cyber-security capability that grew faster than the company anticipated.

And the question this blog exists to answer: GLM-5.3 was officially under review for our pools as of launch day — and went live on August 30 (see the update above). We already serve GLM 5.2 in the Frontier Pool, so 5.3 enters the pipeline as the natural upgrade candidate for that slot. What gates the decision is not quality signals — it’s that the model is API-only today: open weights are promised roughly two weeks out, after Z.ai completes its own safety evaluation. Live status is always on our models-under-review page.

Base modelSame 743B base as GLM 5.2 — not a new pretrain; gains from extended post-training
Context windowZ.ai advertises a 1M-token variant (glm-5.3[1m], with context compaction); the standard-API spec is not yet published
ReasoningEffort levels low / high / max — default max; thinking cannot be disabled on the direct API
Open weightsPublished August 28, 2026zai-org/GLM-5.3 (fp8, 141 shards); at launch they were promised ~2 weeks out
LicenseGLM-5.3 License (custom): commercial use with attribution; Model-as-a-Service operators above US$10B revenue over 12 months must pass Z.ai’s security review. Not MIT — GLM 5.2’s terms did not carry over
List price (API)$1.40 in / $4.40 out per 1M, cached input $0.26 — unchanged from GLM 5.2
AvailabilityFirst-party API access from Z.ai; open weights since August 28; served unlimited on CheapestInference’s Frontier Pool since August 30 — works with Claude Code, OpenCode, Cline and Codex via compatible endpoints

GLM-5.3 benchmarks: what post-training bought

Section titled “GLM-5.3 benchmarks: what post-training bought”

All numbers below are Z.ai’s own reported results — vendor-run, not yet independently reproduced, and the Artificial Analysis index hasn’t rated GLM-5.3 yet. With that caveat on the table, the GLM 5.2 → GLM-5.3 deltas are the story, because the base model is identical:

BenchmarkGLM 5.2GLM-5.3Δ
Terminal-Bench 2.181.088.2+9%
Terminal-Bench 3.04.628.3+515%
DeepSWE v1.146.266.9+45%
SWE-Marathon v1.119.442.5+119%
FrontierSWE67.578.1+16%
NL2Repo48.958.0+19%
Toolathlon Verified59.973.0+22%
AutomationBench v1.0.626.248.2+84%
CyberGym77.284.5+9%

Two readings. The charitable one: the biggest jumps land on the newest, hardest agentic benchmarks (Terminal-Bench 3.0, SWE-Marathon) — exactly where post-training on agent trajectories should show up, and exactly the workloads coding agents run all day. The skeptical one: several of these benchmarks are new or Z.ai-adjacent, and until independent runs land, “strongest open-weights coding model” is a claim, not a fact. Both readings can wait two weeks — the open-weights release is when independent verification becomes possible.

The cyber-defense angle — and why the weights are two weeks out

Section titled “The cyber-defense angle — and why the weights are two weeks out”

The unusual part of this launch is that Z.ai leads with cyber security as a first-class capability, not a footnote. It reports GLM-5.3 at 84.5 on CyberGym — above its figures for Claude Mythos 5 (83.8) and GPT-5.6 Sol (83.6) — and says the model found thousands of real vulnerabilities across open-source projects during training. Z.ai’s framing is defensive: vulnerability detection at scale.

That capability is also the stated reason the weights aren’t out yet. Rather than shipping weights on day one — as it did with GLM 5.2 — Z.ai is running a staged release: API first, then open weights roughly two weeks after launch, once its own safety evaluation and hardening work is complete. Whatever you think of the trade-off, it’s a more deliberate open-weights process than the ecosystem norm, and it puts a concrete clock on the one thing our review is waiting for.

GLM-5.3 vs GLM 5.2 — the model we serve today

Section titled “GLM-5.3 vs GLM 5.2 — the model we serve today”
GLM-5.3GLM 5.2
Base743B (same base)743B
What’s newScaled post-training: agentic coding, tool use, cyber
ContextZ.ai advertises a 1M-token variant (glm-5.3[1m], with context compaction); the standard-API spec is not yet published1M per the model card
ReasoningEffort low / high / max, thinking always onStandard GLM 5.2 semantics
Open weightsPublished August 28 (GLM-5.3 License)Published (MIT)
List price (per 1M)$1.40 in / $4.40 out$1.40 in / $4.40 out
Status hereLive in the Frontier Pool since August 30Retired August 30 — migration

Because the base is unchanged, this isn’t a “new model vs old model” decision so much as a post-training upgrade — the same shape as DeepSeek’s V4-Flash-0731 build, which we upgraded in place in the Core Pool within days of release. If GLM-5.3’s weights land with a usable license and it passes our quality evaluation on real coding and agent workloads, the natural outcome is the same: the Frontier Pool’s GLM slot upgrades, and every existing subscription simply gets the better model.

So the honest status board:

  • Quality — vendor numbers are strong; our own evaluation on real agent workloads (both OpenAI and Anthropic endpoints, tool calling included): ✅ passed.
  • Fit — a post-training upgrade of a model already serving Frontier Pool workloads: as clean as fit gets.
  • Licensing — ✅ open weights published August 28 under the GLM-5.3 License.

Outcome (August 30): live. The upgrade is recorded in the changelog; the models-under-review page is where the next candidate will show up.

Is there an unlimited GLM-5.3 API? Yes — since August 30, 2026 CheapestInference serves GLM-5.3 in the Frontier Pool on flat-rate time-block subscriptions from $71/mo: no token caps during your reserved hours, model id glm-5.3, OpenAI- and Anthropic-compatible.

Does GLM-5.3 have open weights? Yes, since August 28, 2026 — zai-org/GLM-5.3 on Hugging Face, under Z.ai’s custom GLM-5.3 License: commercial use is allowed with attribution, and only Model-as-a-Service operators above US$10B aggregate revenue over 12 months must pass Z.ai’s security review first. GLM 5.2 was MIT; the terms did not carry over.

How much does the GLM-5.3 API cost? Per token, Z.ai lists $1.40 in / $4.40 out per 1M (cached input $0.26) — the same as GLM 5.2. On CheapestInference it is a flat monthly fee: from $71/mo for a daily 8-hour block, unlimited tokens.

What is the difference between GLM-5.3 and GLM 5.2? Same 743B base model — GLM-5.3 is extended post-training on top of it, targeting agentic coding, tool use, and cyber-security workloads. Z.ai reports large gains on agentic benchmarks (SWE-Marathon 19.4 → 42.5, Terminal-Bench 3.0 4.6 → 28.3); all numbers are vendor-run so far.


CheapestInference serves Kimi K3 and Qwen3.8 Max (Flagship Pool), GLM 5.3 and MiniMax M3 (Frontier Pool) and DeepSeek V4.1 Flash and MiMo v2.5 (Core Pool) through one OpenAI- and Anthropic-compatible API on unlimited time-block subscriptions. See the pools or get started.