DeepSeek V4.1 Flash

New
Beta

Agentic coding with image input and 1M context

Chat
Vision
Thinking
JSON Mode
Tool Calling

Pricing

Run instantly. Pay only for what you use.

Input
$0.20/ M
Output
$0.60/ M
Cache read
$0.005/ M

About DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is DeepSeek's multimodal successor to V4 Flash, released under MIT: a 552B-parameter MoE with a separate 196B Engram conditional-memory pool, activating 16B per token while generating (8B during prefill), with a 1M-token context window and native image input through its DeepSeek-ViT encoder.

Compressed Sparse Attention 2 and SWA Bounded Replay cut the global KV cache to 890 bytes per token — roughly a quarter of DeepSeek V4 Flash — so whole-repo prompts and long agent sessions stay fast and affordable at full 1M context.

The result is a clear step up in agentic coding at Flash cost: 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1 at maximum reasoning effort, up from 82.7 and 54.4 for V4 Flash, with function calling, JSON output, and a 1–100 reasoning-effort dial — including none to skip thinking — to pay for depth only where a request needs it.

1Mtoken context window

Sparse attention keeps the global KV cache at 890 bytes per token — about a quarter of DeepSeek V4 Flash.

90.6on Terminal-Bench 2.1

Up from 82.7 for DeepSeek V4 Flash at maximum reasoning effort — from MIT-licensed open weights.

96%avg. cache hit rate on Entrim

Measured on coding-agent workloads — repeated context bills at the cache-read rate.

Key capabilities

DeepSeek V4.1 Flash is strongest where agents reason over a lot of context and now look at screens too: coding loops, tool chains, and document pipelines that need long memory at Flash cost.

Agentic coding

Holds a plan across long tool-call chains — 74.2 on DeepSWE 1.1, up from 54.4 for V4 Flash, and a 3471 Codeforces rating.

Native image input

A from-scratch DeepSeek-ViT encoder reads screenshots, diagrams, and scanned pages directly — 95.6 on DocVQA for the base model — with no separate vision model in the loop.

Reasoning effort from 1 to 100

An integer from 1 to 100 (with low, high, and max presets) sets thinking depth per request, and effort none skips reasoning entirely — pay for depth only where a request needs it.

Reliable structured output

Strict function calling and JSON mode at every reasoning effort — tool arguments and typed pipelines that parse the first time.

Where this model fits

Best for
  • Long-running coding agents

    Plans, edits, and verifies across long tool-call chains — 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1 at maximum reasoning effort.

  • Agents grounded in screenshots and diagrams

    Native image input reads UI screenshots, charts, and scanned pages straight from the prompt, with no separate vision model in the loop.

  • Whole-repo and long-document analysis

    The 1M-token window holds a full repository or document set in one prompt, and sparse attention keeps full-context requests affordable.

  • Browser and workflow automation

    54.8 on AutomationBench — up from 37.7 for V4 Flash — for multi-step tool use across real applications.

Avoid for
  • Fact recall without retrieval

    Standalone knowledge recall is still limited for a sparse-activation model (42.3 on SimpleQA-Verified for the base model) — pair it with RAG for factual workloads.

  • Frontier-grade open-ended reasoning

    Open-ended exam-style reasoning is not the strength — 36.8 on Humanity’s Last Exam without tools. Give research-style questions retrieval and tools, where it reaches 63.9 on HLE.

  • Audio or video inputs

    Inputs are text and images only. Use GLM 5.3 Flash for video, and transcribe audio upstream before sending it to the model.

Performance and
benchmarks

Entrim performance for DeepSeek V4.1 Flash under production load over the last 7 days (daily medians, UTC dates).

Throughput (tokens / sec)

median 131.3tokens/s

Time to first token (TTFT)

median 1,247ms

Public model benchmarks

90.6Terminal-Bench 2.1

Agentic coding and shell tasks run end-to-end in a real terminal.

74.2DeepSWE 1.1

Resolves real software-engineering issues across long tool-call chains.

54.8AutomationBench

Tool-driven automation across real applications and workflows.

Run DeepSeek V4.1 Flash in minutes

Test this model with $10 free credit. Use Entrim's OpenAI-compatible API to call DeepSeek V4.1 Flash, compare output quality, latency, and request cost, then decide if it fits your workload.

Code examples

curl -X POST "https://api.entrim.ai/v1/chat/completions" \  -H "Content-Type: application/json" \  -H "Authorization: Bearer $ENTRIM_API_KEY" \  -d '{    "model": "deepseek-ai/DeepSeek-V4.1-Flash",    "messages": [      {        "role": "user",        "content": "Why do they call it a building if it is already built?"      }    ]  }'

SDKs & docs

Use familiar SDK patterns with quickstart examples for common production setups.

OpenAI-compatible

Keep the request format your team already knows.

No cold starts

Production requests are served without model spin-up delays.

Not sure this is the right model?

Explore related models that may be a better fit for your workload.

High-throughput agentic work with 1M context

1M
FP8
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Multimodal coding agents with 1M context

1M
NVFP4
Vision
Video
Thinking
Input / M
$0.10
Output / M
$0.35
Cached / M
$0.0175

Dense multimodal coding with thinking control

262K
FP8
Vision
Video
Thinking
Input / M
$0.10
Output / M
$0.40
Cached / M
$0.04

Run DeepSeek V4.1 Flash with the Most Cost-Effective LLM Inference

Use $10 free credits to test prompts, token sizes, latency, throughput, output quality, and request cost before moving traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policy•Terms of service