Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

Z.ai just released GLM-5.3. GLM-5.3 runs on the same 743B base model as GLM-5.2. Every reported gain comes from scaled post-training: more task environments, more environment types, longer training. The results land in two places. Coding jumps most on the longest-horizon benchmarks, with Terminal-Bench 3.0 moving from 4.6 to 28.3. Cybersecurity moved further than Z.ai says it expected, with CyberGym reaching 84.5%. Weights are not public yet.

Is It Deployable?

Partially, GLM-5.3 is live through the Z.ai API, the GLM Coding Plan, and ZCode. Weights are not out. Z.ai says it will publish them roughly two weeks after launch, once safety evaluation and hardening finish.

Coding Results

Terminal-Bench 3.0 moves from 4.6 to 28.3 against GLM-5.2. DeepSWE v1.1 moves from 46.2 to 66.9. Agents’ Last Exam (CLI) moves from 23.8 to 28.5. On GDPval-AA v2, which spans 44 occupations, GLM-5.3 scores 1,769.

On Z.ai Code Bench, an internal evaluation, the company reports a 50% improvement over GLM-5.2. It reports 31.4% at roughly 50,000 output tokens per task. Claude Opus 4.8 scores 29.5% at 120,000 tokens. Claude Fable 5 still leads at 39.5% at maximum effort. Z.ai argues a private benchmark reduces contamination risk.

On public suites, GLM-5.3 trails GPT-5.6 Sol and Fable 5 on several harder coding evaluations. All figures are vendor-reported, with harness, context length, and sampling settings documented in the announcement.

The Cybersecurity Result

Z.ai flags this one as unplanned. It added vulnerability-discovery data expecting better single-bug reasoning. Instead, capability kept compounding as training scaled. The model began forming coherent plans across complete exploitation chains.

CyberGym, which tests discovery and validation from white-box source, moves from 77.2% to 84.5%. That edges past Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%. ExploitBench, which requires root-cause reasoning and a working exploit, moves from 24.4% to 54.4%. Mythos 5 sits at 78.0%. On ExploitGym, GLM-5.3 completes 105 tasks in two hours and 130 in six. GLM-5.2 completes 29 and 39. Mythos 5 completes 181 and 247.

The pattern is consistent. The deeper into the exploitation chain a benchmark sits, the larger the gain over GLM-5.2. The gap to closed frontier models also widens.

Interactive Explainer

Key Takeaways


Check out the Z.ai GLM-5.3 technical blog, Zai_org announcement, Z.ai Security Disclosure Ledger and zai-org/GLM-5 on GitHub. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

The post Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPost.

Exit mobile version