Throughput (tokens / sec)
avg. 71.8tokens/s
Frontier-parity coding agents with 1M context
Run instantly. Pay only for what you use.
GLM 5.2 is Z.ai's open-weight engineering flagship, released under MIT: a 744B-parameter MoE with 40B active per token, a 1M-token context window, and up to 128K output tokens, with reasoning effort selectable per request up to a maximum-thinking mode.
It is a step change in long-horizon agentic work: Terminal-Bench 2.1 jumps 62.0 → 81.0 over GLM 5.1, and a context window 5× larger holds entire projects — runs keep engineering standards and project context intact across a full development workflow.
It posts 62.1 on SWE-bench Pro — ahead of GPT-5.5 at 58.6 — and 77.0 on MCP-Atlas for tool-driven work, and its score of 53 on the Artificial Analysis Intelligence Index (ranked #3 of 104 models) is the highest of any model in the Entrim catalog, with strict function calling and JSON output.
62.1on SWE-bench Pro
Ahead of GPT-5.5 (58.6) — the strongest open-weight result on the hardest coding set.
1Mtoken context window
5× the GLM 5.1 window — an entire codebase or document set in one prompt.
#3on the Intelligence Index
Scores 53 on Artificial Analysis — ranked above every other model in the Entrim catalog.
GLM 5.2 is strongest where engineering runs are long and the bar is frontier-level — project-scale refactors, tool-driven automation, and client-side work that has to satisfy real users.
62.1 on SWE-bench Pro — past GPT-5.5 on the contamination-resistant set, from MIT-licensed open weights.
Terminal-Bench 2.1 climbs 62.0 → 81.0 over GLM 5.1 — end-to-end terminal work that used to need a closed model.
Tops Design Arena rankings ahead of closed frontier models — client-side and mobile engineering is this release’s called-out step change.
High and max efforts set thinking depth per request — deep reasoning for the hard steps, faster replies where a request doesn’t need it.
Frontier-grade software engineering
62.1 on SWE-bench Pro — past GPT-5.5 on the contamination-resistant set, from MIT-licensed open weights.
Project-scale agent runs
The 1M-token window holds an entire codebase, and long-horizon execution keeps engineering standards stable across the run.
Front-end and client-side engineering
Tops Design Arena rankings — polished UI and mobile work is the step change Z.ai calls out for this release.
Tool-driven automation at the hard end
77.0 on MCP-Atlas and 54.7 on Humanity's Last Exam with tools — dependable tool selection on multi-step work.
Cost-sensitive, high-volume fleets
Premium 40B-active pricing. For always-on loops where per-token cost dominates, use Qwen 3.6 35B-A3B.
Image, audio, or video inputs
GLM 5.2 is text-only. For agentic work grounded in screenshots, use Qwen 3.8 27B.
Tight output-token budgets at high effort
Reasoning traces run long — about 1.4× the median output tokens in Artificial Analysis testing. Use a lower reasoning effort for cost-sensitive replies.
avg. 71.8tokens/s
avg. 720ms
62.1SWE-bench Pro
Contamination-resistant software engineering — ahead of GPT-5.5 at 58.6.
81.0Terminal-Bench 2.1
Agentic coding and shell tasks run end-to-end in a real terminal.
77.0MCP-Atlas
Tool use through the Model Context Protocol across multi-step workflows.
Code examples
curl -X POST "https://api.entrim.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $ENTRIM_API_KEY" \ -d '{ "model": "zai-org/GLM-5.2", "messages": [ { "role": "user", "content": "Why do they call it a building if it is already built?" } ] }'SDKs & docs
Use familiar SDK patterns with quickstart examples for common production setups.
OpenAI-compatible
Keep the request format your team already knows.
No cold starts
Production requests are served without model spin-up delays.
Multimodal coding agents with 1M context
High-throughput agentic work with 1M context
Dense multimodal coding with thinking control