Gemma 4 26B A4B

Gemma 4 quality at 4B-active speed

Chat
Multi-Modal
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.05/ M
Output
$0.25/ M
Cache read
$0.025/ M

About Gemma 4 26B A4B

Gemma 4 26B A4B is the sparse MoE member of Google DeepMind's Gemma 4 family, released under Apache 2.0: 25.2B total parameters with just 3.8B active per token across 128 fine-grained experts, a 256K-token context window, and native text + image input, pre-trained across 140+ languages.

It keeps near-flagship quality at small-model compute: within about 3 points of the dense Gemma 4 31B on MMLU-Pro and MMMU-Pro, with a configurable thinking mode that adds step-by-step reasoning only where a request needs it.

With thinking enabled it posts 82.6% on MMLU-Pro, 88.3 on AIME 2026, and 73.8 on MMMU-Pro — the speed-and-cost pick of the Gemma 4 line, with strict function calling and JSON output for production pipelines.

3.8Bactive of 25.2B total

128 fine-grained experts route per token — near-31B quality at small-model compute.

256Ktoken context window

Holds long documents, retrieved passages, and multi-turn sessions in one prompt.

140+languages in pre-training

35+ supported out of the box — one deployment serves a global user base.

Key capabilities

Gemma 4 26B A4B is strongest where volume meets variety — multilingual assistants, visual document pipelines, and high-throughput extraction that need flagship-adjacent quality without flagship cost.

Near-flagship quality

82.6% on MMLU-Pro and an Arena ELO within a dozen points of Gemma 4 31B — from a fraction of the compute per token.

Native vision

73.8 on MMMU-Pro — screenshots, PDFs, and diagrams are read alongside text, no separate vision model to route to.

Configurable thinking

Step-by-step reasoning on demand — 88.3 on AIME 2026 with thinking enabled, fast direct replies with it off.

Reliable structured output

Strict function calling and JSON mode — tool arguments and extraction pipelines that parse the first time.

Where this model fits

Best for
  • Multilingual chat and support at scale

    35+ languages out of the box, 140+ in pre-training — one deployment serves a global user base at 4B-active cost.

  • Multimodal document and screenshot pipelines

    Native vision at 73.8 on MMMU-Pro reads screenshots, PDFs, and diagrams alongside text in one model.

  • High-volume fleets on a small-model budget

    Arena-ranked within a dozen ELO points of the 31B flagship while activating 3.8B parameters per token.

  • Document and image extraction to JSON

    Function calling and JSON mode return schema-valid output for ticket, form, and invoice pipelines.

Avoid for
  • The hardest coding and reasoning ceiling

    The dense Gemma 4 31B scores higher across the board — step up to it when the last few points matter.

  • Tight output-token budgets with thinking on

    Thinking mode is verbose — about 2× the median output tokens in Artificial Analysis testing. Cap or disable thinking for cost-sensitive replies.

  • Audio input or speech understanding

    Inputs are text and image only. Transcribe audio upstream before sending it to the model.

Performance and
benchmarks

Entrim performance for Gemma 4 26B A4B under production load in the past few days.

Throughput (tokens / sec)

avg. 215tokens/s

Time to first token (TTFT)

avg. 540ms

Public model benchmarks

82.6%MMLU-Pro

Graduate-level knowledge and reasoning across 14 domains.

73.8MMMU-Pro

University-level reasoning over images, charts, and diagrams.

88.3AIME 2026

Competition mathematics measuring multi-step reasoning.

Run Gemma 4 26B A4B in minutes

Test this model with $25 free credit. Use Entrim's OpenAI-compatible API to call Gemma 4 26B A4B, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "google/gemma-4-26B-A4B-it",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

Multimodal long-context work with thinking control

256K
FP8
Vision
Multilingual
Input / M
$0.10
Output / M
$0.30
Cached / M
$0.04

Tool-heavy coding at 3B-active cost

262K
FP8
Vision
Multimodal
Video
Input / M
$0.05
Output / M
$0.25
Cached / M
$0.025

Low-latency reasoning on a tiny budget

131.1K
FP4
Thinking
Input / M
$0.02
Output / M
$0.10
Cached / M
$0.005

Run Gemma 4 26B A4B with the Most Cost-Effective LLM Inference

Use $25 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service