DeepSeek V4 Flash

Version 0731

High-throughput agentic work with 1M context

Chat
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.09/ M
Output
$0.17/ M
Cache read
$0.015/ M

About DeepSeek V4 Flash

DeepSeek V4 Flash is DeepSeek's efficiency-focused open-weight model: a 284B-parameter MoE that activates just 13B per token, with a 1M-token context window and MIT-licensed weights. Three thinking levels — non-think, think-high, and think-max — let you pick speed or reasoning depth per request.

Its hybrid sparse attention is built for long context: at 1M tokens it runs on roughly 27% of the compute and 10% of the KV cache of DeepSeek V3.2, so whole-repo prompts and heavily cached agent context stay fast and affordable.

The result is production-grade agentic coding at a small-model price: 79% on SWE-bench Verified and 91.6 on LiveCodeBench in max-thinking mode, with strict function calling and JSON output available in every mode.

1Mtoken context window

Sparse attention runs 1M-token prompts on ~27% of the compute of DeepSeek V3.2.

13Bactive of 284B total

MoE routing keeps per-request cost and latency near small-model levels.

3thinking levels

Non-think, think-high, and think-max trade latency for reasoning depth per request.

Key capabilities

DeepSeek V4 Flash is strongest where agents reason over a lot of context: coding loops, tool chains, and document pipelines that need long memory without dense-model cost.

Agentic coding

Holds a plan across long tool-call chains — the 0731 re-post-train pushes Terminal-Bench 2.1 to 82.7, up from 61.8 on the preview build.

Three thinking levels

Non-think, think-high, and think-max — the 0731 re-post-train sharpens the reasoning modes hardest (DeepSWE 7.3 → 54.4), so you pay for depth only where a request needs it.

1M-token context

Hybrid sparse attention reads 1M tokens with ~10% of the KV cache of V3.2 — whole repos and document sets in one prompt.

Reliable structured output

Strict function calling and JSON mode in every thinking level — tool arguments and typed pipelines that parse the first time.

Where this model fits

Best for
  • Long-running coding agents

    Plans, edits, and verifies across long tool-call chains — 79% on SWE-bench Verified in max-thinking mode.

  • Whole-repo and long-document analysis

    The 1M-token window holds a full repository or document set in one prompt, and sparse attention keeps it affordable.

  • High-volume, cost-sensitive fleets

    13B active parameters and deep cache-read discounts keep per-request cost flat as traffic scales.

  • Structured extraction to JSON

    Strict function calling and JSON mode return schema-valid output for pipelines — in every thinking mode.

Avoid for
  • Image, audio, or video inputs

    DeepSeek V4 Flash is text-only. Route visual inputs to a multimodal model such as Gemma 4 31B or Qwen 3.6 35B-A3B.

  • Fact recall without retrieval

    Like other small-activation models it is weak on standalone knowledge recall (34.1% on SimpleQA-Verified) — pair it with RAG for factual workloads.

  • Tight output-token budgets in think-max

    Max thinking emits long reasoning traces that drive up output tokens and latency. Use non-think or think-high for cost- and latency-sensitive replies.

Estimate your Savings

Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.

Performance and
benchmarks

Entrim performance for DeepSeek V4 Flash under production load in the past few days.

Throughput (tokens / sec)

avg. 150tokens/s

Time to first token (TTFT)

avg. 820ms

Public model benchmarks

82.7Terminal-Bench 2.1

Agentic coding and shell tasks run end-to-end in a real terminal.

54.4DeepSWE

Resolves real software-engineering issues across long tool-call chains.

70.3Toolathlon (verified)

Sustains long multi-tool workflows without losing the plan.

Run DeepSeek V4 Flash in minutes

Test this model with $25 free credit. Use Entrim's OpenAI-compatible API to call DeepSeek V4 Flash, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "deepseek-ai/DeepSeek-V4-Flash",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

Tool-heavy coding at 3B-active cost

262K
FP8
Vision
Multimodal
Video
Input / M
$0.05
Output / M
$0.25
Cached / M
$0.025

o4-mini-class reasoning for tool-driven coding

131.1K
FP4
Thinking
Input / M
$0.029
Output / M
$0.13
Cached / M
$0.01

Multimodal long-context work with thinking control

256K
FP8
Vision
Multilingual
Input / M
$0.10
Output / M
$0.30
Cached / M
$0.04

Run DeepSeek V4 Flash with the Most Cost-Effective LLM Inference

Use $25 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service