Zhipu AI released GLM-5.3 on August 14, 2026, and the headline number isn't the parameter count. It's the absence of a pretraining run. GLM-5.3 shares the exact same 743B base model as GLM-5.2 — same weights, same architecture, same token budget already spent. Every gain in coding, agentic tool use, and the surprise cybersecurity capability came from what Zhipu calls "post-training Scaling": dozens of× more long-horizon task environments, richer environment types, and drastically longer RL post-training on top of a frozen base.
The result is the strongest open-weights coding model Zhipu has ever shipped, and a clean experimental answer to a question the whole industry has been circling: how much farther can you push a frozen base model without spending another $100M on pretraining? The answer, at least this month, is further than most people expected. Terminal-Bench 3.0 went from 4.6 to 28.3. DeepSWE v1.1 went from 46.2 to 66.9. And in white-box vulnerability reasoning, GLM-5.3 scored 84.5% on CyberGym — above Mythos 5's 83.8% and GPT-5.6 Sol's 83.6%.
What's New: Same Base, +50% Coding Feel

GLM-5.3 is not a new foundation. It is GLM-5.2 with a radically expanded post-training regime, and Zhipu is explicit about that. The four things that actually changed:
Coding is the lead event. On Zhipu's internal Z.ai Code Bench — a real local-dev-environment evaluation that runs end-to-end tasks at different thinking tiers — GLM-5.3 improved 50% over GLM-5.2. At the High effort tier it hit 31.4% accuracy, above Claude Opus 4.8's best-tier 29.5%, while spending ~50K output tokens per task versus Opus 4.8's ~120K. That token-efficiency gap is the number that matters for production agent loops: same task, less than half the output budget.
Public benchmarks that matter moved hard. Terminal-Bench 3.0 — real terminal-task completion — jumped from 4.6 to 28.3, open-source best. DeepSWE v1.1, which measures long-horizon software engineering and sustained code modification, went 46.2 → 66.9. Agents' Last Exam (CLI) went 23.8 → 28.5. AutomationBench went 26.2 → 48.2. Toolathlon Verified went 59.9 → 73.0. GDPval-AA v2, which covers 44 professions, went 1508 → 1769 Elo.
Cybersecurity emerged as an unplanned capability. This is the one Zhipu itself calls out as surprising. Security work, they argue, is essentially "strictly constrained coding" — and as they expanded the post-training task environments to longer, more constrained expert workflows, the model started doing serious security work without being explicitly trained to.
Open weights in two weeks. GLM-5.3 went live immediately on ZCode, AutoClaw, the full GLM Coding Plan tier, and third-party coding platforms (TraeWork/TraeCode, Coze, WorkBuddy/CodeBuddy, Qoder, CatPaw, JoyCode, OpenCode). The API followed on August 19. Full weights landed on HuggingFace roughly two weeks later, after security hardening — with a "trusted access" tier for the most sensitive offensive-security capabilities.
Post-Training Scaling: The Textbook Didn't Change

"Post-training scaling" is the idea driving this release, and it's worth defining precisely because the industry has been sloppy about the term. It means: pretraining is finished and frozen. You do not touch the base weights. You raise the model's capability ceiling entirely through post-training — RL, data quality, reward design, and the scale of the environments you train on. Zhipu's own analogy: the textbook didn't change, but the training method got much better, so the student's grades went up.
The infrastructure stack that made this concrete is three pieces:
- IndexShare — long-context post-training infrastructure that keeps long horizon tasks tractable.
- SAO — long-horizon task RL, extending training far beyond short coding problems.
- Slime — a new-generation large-scale asynchronous training framework for running these environments at scale.
What does "dozens of× more long-horizon environments" look like in practice? Zhipu describes tasks whose workload is "equivalent to an engineer working for several days straight." For machine-learning optimization tasks, the model uses the same compute cluster, storage, internal docs, codebases, and experiment results a real algorithm engineer would use, and must deliver a quantifiable speedup end-to-end. The model isn't just solving coding problems — it's learning to find the problem, scope the analysis, implement the fix, and run validation.
That's why the cybersecurity numbers matter as much as the coding ones. Zhipu says they started investing in this direction in September 2025, working with top domestic security labs to build longer task environments and specialized evaluation harnesses. As the post-training scale grew, security capability "emerged" — it wasn't a targeted SFT push. In CyberGym (white-box source code → trigger faults → identify and verify vulnerabilities), GLM-5.3 went from 77.2% to 84.5%. In ExploitBench (understanding root cause and writing working exploits), it more than doubled from 24.4% to 54.4% — though it still trails Mythos 5 (78.0%) and GPT-5.6 Sol (76.5%) on the harder exploitation side. In ExploitGym (throughput-normalized exploit completion), it finished 105 tasks in 2 hours and 130 in 6 hours, versus 29/39 for GLM-5.2 and 181/247 for Mythos 5.
The pre-release red-team process is the detail that tells you how seriously they take this. Over two weeks, working with Tsinghua, Nankai, and domestic security teams, they ran GLM-5.3 across 269 projects and found 2,436 valid vulnerabilities — 1,088 of them medium-to-high severity — with the oldest bug tracing back roughly 40 years. Vulnerabilities were found in Android, Windows, macOS, and apps people use daily. Zhipu also launched an "Open Source Shield" program to offer continuous security audits to open-source projects. Their logic for open-sourcing the offensive capability is blunt: if strong attack capability is already spreading through the ecosystem, defense capability cannot live inside a handful of closed companies.
Benchmarks: Open-Source Best, Approaching the Closed Frontier

The comparison set in Zhipu's own charts is worth reading closely because it tells you who they consider the competition: GLM-5.3, GLM-5.2, Kimi K3, Fable 5, and GPT-5.6 Sol.
| Benchmark | GLM-5.3 | GLM-5.2 | Kimi K3 | Fable 5 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 3.0 | 28.3 | 4.6 | 17.4 | 33.7 | 34.6 |
| DeepSWE v1.1 | 66.9 | 46.2 | 67.5 | 69.7 | 72.7 |
| Agents' Last Exam (CLI) | 28.5 | 23.8 | 27.6 | 23.8 | 28.6 |
| AutomationBench | 48.2 | 26.2 | 46.7 | 46.2 | 45.8 |
| HLE w/ Tools | 62.5 | 54.7 | 59.8 | 63.9 | 64.5 |
| GDPval-AA v2 (Elo) | 1769 | 1508 | 1682 | 1743 | 1730 |
A few readings:
GLM-5.3 is the open-source leader on Terminal-Bench 3.0, Agents' Last Exam, AutomationBench, and GDPval-AA v2. It ties or edges Kimi K3 on nearly every agentic benchmark. On DeepSWE it's 66.9 against Kimi K3's 67.5 — essentially tied, and only a few points behind the closed frontier (Fable 5 at 69.7, GPT-5.6 Sol at 72.7).
The caution flags are real. These are all Zhipu self-reported numbers; no independent third-party reproduction exists yet. The gap to Fable 5 and GPT-5.6 Sol on Terminal-Bench and DeepSWE is still 5–6 points. And on ExploitBench, the hardest security benchmark, GLM-5.3's 54.4% is well behind Mythos 5's 78.0%. The defensive/code-review half of security is where it's strong; active exploitation is not there yet.
Pricing positions GLM-5.3 as a flagship, not a cheap model: roughly ¥10/M input and ¥31/M output (~$1.40 / ~$4.40) at the API level. That's above DeepSeek V4-Pro and well above any Flash-tier model. You are paying for top-tier open coding and agent capability, not throughput.
The bigger industry frame is the monthly cadence. July 17: Moonshot released Kimi K3 at 2.8T parameters, the first open model past the 3T scale. August 13: DeepSeek V4-Pro official went live at 1.6T total / 49B active. August 14: GLM-5.3. In four weeks, three of China's top labs shipped flagship iterations, and all three pushed coding toward the international frontier. The competition has moved from single-metric races to a fight over high-value scenarios — coding agents, security auditing, enterprise work.
What It Means for Developers

The practical decision space for builders:
Self-hosting vs API. GLM-5.3 weights are open on HuggingFace, so you can run it. But 743B total parameters is not a workstation model — plan on serious multi-GPU or quantized deployment. The API path (Z.ai direct, OpenRouter, or China's National Supercomputing Internet service) is the realistic route for most teams.
When to pick GLM-5.3 over GLM-5.2. If your workload is long-horizon agentic coding, terminal tasks, or real software-engineering loops, the 4.6 → 28.3 Terminal-Bench jump and the 50% internal coding-feel improvement are the argument. The token efficiency on Z.ai Code Bench (~50K vs Opus 4.8's ~120K per task) directly translates to lower API bills on long agent runs.
Security-audit use cases. For code review, vulnerability discovery, and defensive security work, GLM-5.3 is now the open model to test. The 84.5% CyberGym score and the 2,436-vulnerability red-team run are real signal. For exploit development and offensive benchmarking, it's not yet at Mythos 5 level.
Routing between GLM-5.3 and the Flash tier. When GLM-5.3-Flash ships two weeks later at 1/10 the price, the split is straightforward: GLM-5.3 for the hardest agentic coding and security work; Flash for high-volume multimodal throughput.
The strategic bet to watch. If post-training scaling on a frozen base keeps producing these gains, the economic logic of the whole industry changes. Pretraining is the capital-intensive part. If you can keep raising the ceiling through post-training — better environments, longer RL, richer data — then the lab that already owns a strong base model has a durable advantage without repeatedly spending nine figures on new pretraining runs. Zhipu's own line: they may still be far from exhausting their base model's intelligence ceiling.
Bottom Line
GLM-5.3 is a proof point more than a product. Same 743B base as GLM-5.2, zero new pretraining, and yet the strongest open coding model in the world on Terminal-Bench 3.0, a doubled cybersecurity score on CyberGym, and a 50% internal coding-feel gain over the predecessor. The post-training-scaling recipe — IndexShare, SAO, the Slime framework, and environments that approximate days of expert engineering work — is the actual release. Weights open-sourced within two weeks, API live at flagship pricing. For developers, it's the new default open model to benchmark your coding agent against; for the industry, it's evidence that the next round of capability gains may come from training methodology, not from bigger base models.
Resources
- Z.ai Research Blog: GLM-5.3 — official technical announcement and benchmark breakdown
- Z.ai Technical Blog — English-language deep dive on post-training scaling
- HuggingFace: GLM-5.3 Weights — open weights released ~2 weeks post-launch
- Securities Times Report — launch coverage and the one-month domestic iteration context
- The Decoder Coverage — English analysis of open-source coding claims
GLM-5.3 is available via API at z.ai and on OpenRouter. Weights are open-sourced on HuggingFace under Zhipu's standard open-weight terms.