Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context

Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series and the cheapest capable coding model the lab has shipped. It is a mixture-of-experts model with 320B total parameters and 18B active per token, a 1,048,576-token context window, and image and video input — released under an MIT license with weights on Hugging Face. According to Z.ai reports, it beats GLM-5.2 across benchmarks and real workloads at roughly one-tenth the price, while landing within half a point of Claude Opus 4.8 on its internal coding benchmark. The model spent its first week running anonymously as “Ox Alpha” on OpenCode and OpenRouter, served entirely on domestically produced Chinese AI chips.

Is it deployable?

Yes, on two tracks. The weights are live on Hugging Face under an MIT license, and a hosted API is already priced and serving.

The architecture is where the efficiency comes from

GLM-5.3-Flash starts from a newly trained base model on a 30T-token multimodal corpus. Three changes are worth knowing:

Benchmarks

Most numbers mentioned in the table below are Z.ai-reported and the harnesses differ per test — the model card’s footnotes specify temperature, context limits and judge models per benchmark, so treat cross-model comparisons as setup-dependent.

Benchmark GLM-5.3-Flash Reference
Terminal-Bench 2.1 84.3 Opus 4.8: 85.0 · GPT-5.6 Terra: 87.4
DeepSWE v1.1 63.4 GLM-5.2: 46.2
AutomationBench 48.8 GLM-5.2: 26.2
HLE 55.3
OfficeQA Pro 62.4 ahead of Opus 4.8
Z.ai Code Bench v1.0 (max) 29.0 Opus 4.8: 29.5

Independently, Artificial Analysis scores it 57 on the Intelligence Index, with 48.7 output tokens/sec and 1.52s TTFT on Z.ai’s API — strong intelligence-per-dollar, but slow and verbose. Vision is the weak flank: it trails Gemini 3.7 Flash on BabyVision and MVbench.

The serving story is the underreported part

Z.ai states the entire Ox Alpha preview ran on domestically produced Chinese AI chips, using a custom SGLang-based engine that disaggregates encoding, prefill and decoding, and reports a 3× end-to-end serving improvement across tens of thousands of accelerators.

Pricing and access

Standard API pricing is $0.15/M input, $0.03/M cached input, $0.50/M output. Z.ai reports a score of 57 on Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task on the discounted tier. The model is live for all GLM Coding Plan tiers — Lite ($18/mo), Pro ($80), Max ($168) — at 3× the usable quota of GLM-5.3, and its multimodal capabilities surface in ZCode through Browser Use and Computer Use. Local serving is supported on SGLang, vLLM, TokenSpeed and KTransformers.

Key Takeaways


Check out the Model Weights and Blog. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us

The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost.

Exit mobile version