GPT OSS 120B

o4-mini-class reasoning for tool-driven coding

Chat
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.029/ M
Output
$0.13/ M
Cache read
$0.01/ M

About GPT OSS 120B

GPT OSS 120B is OpenAI's most capable open-weight model, released under Apache 2.0: a 117B-parameter MoE with 5.1B active per token and a 131K-token context window, trained on the Harmony response format with configurable reasoning effort — low, medium, or high.

The effort dial is the point: AIME 2024 climbs from 56.3% at low to 95.8% at high effort, so one deployment serves quick routine calls and hard reasoning alike — while 5.1B-active routing keeps it among the fastest open models served.

At high effort it posts 80.1% on GPQA Diamond and 62.4% on SWE-bench Verified, with Harmony-trained function calling and structured JSON output — OpenAI's published evals place it near o4-mini on core reasoning.

5.1Bactive of 117B total

MoE routing makes it one of the fastest open models served at scale.

3reasoning effort levels

Low, medium, and high — AIME 2024 jumps from 56.3% to 95.8% as effort rises.

131Ktoken context window

Fits long documents and full tool-call transcripts in a single request.

Key capabilities

GPT OSS 120B is strongest where request difficulty varies — reasoning APIs, analysis endpoints, and tool-calling backends that need depth on some calls and speed on the rest.

Configurable reasoning effort

One dial per request — AIME 2024 climbs 56.3 → 80.4 → 95.8 across low, medium, and high effort.

Graduate-level science

80.1% on GPQA Diamond at high effort — analysis that holds up in technical and scientific domains.

Fast tool calling

Top-3 output speed among open models on Artificial Analysis, with Harmony-trained function calling.

Reliable structured output

Strict JSON schemas and tool arguments that parse the first time — built for API backends.

Where this model fits

Best for
  • Reasoning APIs with mixed difficulty

    Per-request effort control — low for the routine 95%, high for the hard 5%. One model, one deployment.

  • Math, science, and technical analysis

    95.8% on AIME 2024 and 80.1% on GPQA Diamond at high effort — competition-grade reasoning on demand.

  • Latency-sensitive tool calling

    Top-3 output speed among open models measured by Artificial Analysis, with Harmony-trained function calling.

  • Structured outputs at volume

    Strict JSON schemas and tool calls that parse the first time, at a 5.1B-active price point.

Avoid for
  • Repo-scale coding agents

    62.4% on SWE-bench Verified trails dedicated coding models — use DeepSeek V4 Flash or Qwen 3.6 27B for issue-to-patch work.

  • Image, audio, or video inputs

    GPT OSS 120B is text-only. Route visual inputs to a multimodal model such as Qwen 3.6 35B-A3B or Gemma 4 31B.

  • Tight output-token budgets at high effort

    High effort produces long reasoning traces — about 1.6× the median output tokens in Artificial Analysis testing. Use low or medium effort for cost-sensitive replies.

Estimate your Savings

Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.

Performance and
benchmarks

Entrim performance for GPT OSS 120B under production load in the past few days.

Throughput (tokens / sec)

avg. 218.4tokens/s

Time to first token (TTFT)

avg. 420ms

Public model benchmarks

95.8%AIME 2024

Competition mathematics measuring multi-step reasoning.

80.1%GPQA Diamond

Graduate-level scientific reasoning and analysis.

62.4%SWE-bench Verified

Real GitHub issues resolved end-to-end in agentic runs.

Run GPT OSS 120B in minutes

Test this model with $25 free credit. Use Entrim's OpenAI-compatible API to call GPT OSS 120B, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "openai/gpt-oss-120b",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

Low-latency reasoning on a tiny budget

131.1K
FP4
Thinking
Input / M
$0.02
Output / M
$0.10
Cached / M
$0.005

DeepSeek V4 Flash

Version 0731

High-throughput agentic work with 1M context

1M
FP4
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Dense agentic coding with thinking control

262K
FP8
Vision
Multimodal
Video
Input / M
$0.10
Output / M
$0.40
Cached / M
$0.04

Run GPT OSS 120B with the Most Cost-Effective LLM Inference

Use $25 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service