GLM 5.2

Frontier-parity coding agents with 1M context

Chat
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.9/ M
Output
$3.0/ M
Cache read
$0.15/ M

About GLM 5.2

GLM 5.2 is Z.ai's open-weight engineering flagship, released under MIT: a 744B-parameter MoE with 40B active per token, a 1M-token context window, and up to 128K output tokens, with reasoning effort selectable per request up to a maximum-thinking mode.

It is a step change in long-horizon agentic work: Terminal-Bench 2.1 jumps 62.0 → 81.0 over GLM 5.1, and a context window 5× larger holds entire projects — runs keep engineering standards and project context intact across a full development workflow.

It posts 62.1 on SWE-bench Pro — ahead of GPT-5.5 at 58.6 — and 77.0 on MCP-Atlas for tool-driven work, and its score of 53 on the Artificial Analysis Intelligence Index (ranked #3 of 104 models) is the highest of any model in the Entrim catalog, with strict function calling and JSON output.

62.1on SWE-bench Pro

Ahead of GPT-5.5 (58.6) — the strongest open-weight result on the hardest coding set.

1Mtoken context window

5× the GLM 5.1 window — an entire codebase or document set in one prompt.

#3on the Intelligence Index

Scores 53 on Artificial Analysis — ranked above every other model in the Entrim catalog.

Key capabilities

GLM 5.2 is strongest where engineering runs are long and the bar is frontier-level — project-scale refactors, tool-driven automation, and client-side work that has to satisfy real users.

Frontier-parity engineering

62.1 on SWE-bench Pro — past GPT-5.5 on the contamination-resistant set, from MIT-licensed open weights.

A generational agentic jump

Terminal-Bench 2.1 climbs 62.0 → 81.0 over GLM 5.1 — end-to-end terminal work that used to need a closed model.

Design-grade front-end output

Tops Design Arena rankings ahead of closed frontier models — client-side and mobile engineering is this release’s called-out step change.

Selectable reasoning effort

High and max efforts set thinking depth per request — deep reasoning for the hard steps, faster replies where a request doesn’t need it.

Where this model fits

Best for
  • Frontier-grade software engineering

    62.1 on SWE-bench Pro — past GPT-5.5 on the contamination-resistant set, from MIT-licensed open weights.

  • Project-scale agent runs

    The 1M-token window holds an entire codebase, and long-horizon execution keeps engineering standards stable across the run.

  • Front-end and client-side engineering

    Tops Design Arena rankings — polished UI and mobile work is the step change Z.ai calls out for this release.

  • Tool-driven automation at the hard end

    77.0 on MCP-Atlas and 54.7 on Humanity's Last Exam with tools — dependable tool selection on multi-step work.

Avoid for
  • Cost-sensitive, high-volume fleets

    Premium 40B-active pricing. For always-on loops where per-token cost dominates, use Qwen 3.6 35B-A3B.

  • Image, audio, or video inputs

    GLM 5.2 is text-only. For agentic work grounded in screenshots, use Qwen 3.8 27B.

  • Tight output-token budgets at high effort

    Reasoning traces run long — about 1.4× the median output tokens in Artificial Analysis testing. Use a lower reasoning effort for cost-sensitive replies.

Performance and
benchmarks

Entrim performance for GLM 5.2 under production load in the past few days.

Throughput (tokens / sec)

avg. 71.8tokens/s

Time to first token (TTFT)

avg. 720ms

Public model benchmarks

62.1SWE-bench Pro

Contamination-resistant software engineering — ahead of GPT-5.5 at 58.6.

81.0Terminal-Bench 2.1

Agentic coding and shell tasks run end-to-end in a real terminal.

77.0MCP-Atlas

Tool use through the Model Context Protocol across multi-step workflows.

Run GLM 5.2 in minutes

Test this model with $10 free credit. Use Entrim's OpenAI-compatible API to call GLM 5.2, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "zai-org/GLM-5.2",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

GLM 5.3 Flash

New
Beta
Turbo

Multimodal coding agents with 1M context

1M
NVFP4
Vision
Multimodal
Video
Input / M
$0.10
Output / M
$0.35
Cached / M
$0.0175

High-throughput agentic work with 1M context

1M
FP8
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Dense multimodal coding with thinking control

262K
FP8
Vision
Multimodal
Video
Input / M
$0.10
Output / M
$0.40
Cached / M
$0.04

Run GLM 5.2 with the Most Cost-Effective LLM Inference

Use $10 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service