Zhipu officially released GLM-5.3 today. The biggest difference from the previous generation is that the base model remains unchanged, and all improvements come from scaling in the post-training phase. Through a tenfold increase in the scale of long-range task environments, more diverse environment types, and extended post-training time, Zhipu has significantly raised the upper limit of the model's intelligence on the same base. Internal company evaluations show that the programming experience of the new model has improved by about 50% compared to GLM-5.2, making it the strongest open-source model in terms of programming ability currently.
On public benchmarks, GLM-5.3 also delivered impressive results. Terminal-Bench 3.0 scores rose from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, Agents' Last Exam from 23.8 to 28.5, and GDPval-AA v2 reached 1769 points. Its programming and agent capabilities are now close to Claude Fable 5, surpassing other domestic models in experience. It has taken first place in the open-source category in both Terminal Bench 3.0 and Agents' Last Exam (CLI) tests.

Post-training Scaling: The Potential of the Base Model Is Far From Being Fully Explored
Zhipu emphasized that all the above improvements come from post-training rather than changing the model. Based on IndexShare, SAO, and the continuously evolving next-generation Slime framework, the team efficiently advanced reinforcement learning on the same base as GLM-5.2, and admitted, "it may be far from fully exploring the intelligence upper limit of this base." This approach sends a clear signal: competition in large models does not always rely on increasing parameters; thoroughly training an existing base can also approach the cutting edge.

In terms of security, GLM-5.3 performs equally well as Mythos 5 in tasks such as white-box code review and vulnerability detection, showing potential for network defense scenarios. Zhipu will release the full model weights two weeks after the release, but only after completing security assessments and model hardening to limit potential attack capabilities and retain defensive value. At the same time, the new model is now available on the official programming tools ZCode, AutoClaw, and GLM Coding Plan, and offers early access to coding platforms such as Trae, Kouzi, WorkBuddy/CodeBuddy, and Qoder. As open-source models approach closed-source flagship models in the programming track, the second half of the competition among domestic large models is shifting from "who speaks better" to "who works better."
