GPT OSS 120B

o4-mini-class reasoning for tool-driven coding

Chat
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.029/ M
Output
$0.13/ M
Cache read
$0.01/ M

About GPT OSS 120B

GPT OSS 120B is OpenAI's most capable open-weight model, released under Apache 2.0: a 117B-parameter MoE with 5.1B active per token and a 131K-token context window, trained on the Harmony response format with configurable reasoning effort — low, medium, or high.

The effort dial is the point: AIME 2024 climbs from 56.3% at low to 95.8% at high effort, so one deployment serves quick routine calls and hard reasoning alike — while 5.1B-active routing keeps it among the fastest open models served.

At high effort it posts 80.1% on GPQA Diamond and 62.4% on SWE-bench Verified, with Harmony-trained function calling and structured JSON output — OpenAI's published evals place it near o4-mini on core reasoning.

5.1Bactive of 117B total

MoE routing makes it one of the fastest open models served at scale.

3reasoning effort levels

Low, medium, and high — AIME 2024 jumps from 56.3% to 95.8% as effort rises.

131Ktoken context window

Fits long documents and full tool-call transcripts in a single request.

Key capabilities

GPT OSS 120B is strongest where request difficulty varies — reasoning APIs, analysis endpoints, and tool-calling backends that need depth on some calls and speed on the rest.

Configurable reasoning effort

One dial per request — AIME 2024 climbs 56.3 → 80.4 → 95.8 across low, medium, and high effort.

Graduate-level science

80.1% on GPQA Diamond at high effort — analysis that holds up in technical and scientific domains.

Fast tool calling

Top-3 output speed among open models on Artificial Analysis, with Harmony-trained function calling.

Reliable structured output

Strict JSON schemas and tool arguments that parse the first time — built for API backends.

Where this model fits

Best for
  • Reasoning APIs with mixed difficulty

    Per-request effort control — low for the routine 95%, high for the hard 5%. One model, one deployment.

  • Math, science, and technical analysis

    95.8% on AIME 2024 and 80.1% on GPQA Diamond at high effort — competition-grade reasoning on demand.

  • Latency-sensitive tool calling

    Top-3 output speed among open models measured by Artificial Analysis, with Harmony-trained function calling.

  • Structured outputs at volume

    Strict JSON schemas and tool calls that parse the first time, at a 5.1B-active price point.

Avoid for
  • Repo-scale coding agents

    62.4% on SWE-bench Verified trails dedicated coding models — use DeepSeek V4 Flash or Qwen 3.8 27B for issue-to-patch work.

  • Image, audio, or video inputs

    GPT OSS 120B is text-only. Route visual inputs to a multimodal model such as Qwen 3.8 27B or Gemma 4 31B.

  • Tight output-token budgets at high effort

    High effort produces long reasoning traces — about 1.6× the median output tokens in Artificial Analysis testing. Use low or medium effort for cost-sensitive replies.

Estimate your Savings

Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.

Benchmarks

Public model benchmarks

95.8%AIME 2024

Competition mathematics measuring multi-step reasoning.

80.1%GPQA Diamond

Graduate-level scientific reasoning and analysis.

62.4%SWE-bench Verified

Real GitHub issues resolved end-to-end in agentic runs.

Run GPT OSS 120B in minutes

Test this model with $10 free credit. Use Entrim's OpenAI-compatible API to call GPT OSS 120B, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "openai/gpt-oss-120b",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

High-throughput agentic work with 1M context

1M
FP8
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Dense multimodal coding with thinking control

262K
FP8
Vision
Video
Thinking
Input / M
$0.10
Output / M
$0.40
Cached / M
$0.04

Run GPT OSS 120B with the Most Cost-Effective LLM Inference

Use $10 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policy•Terms of service