Throughput (tokens / sec)
median 131.3tokens/s
Agentic coding with image input and 1M context
Run instantly. Pay only for what you use.
DeepSeek V4.1 Flash is DeepSeek's multimodal successor to V4 Flash, released under MIT: a 552B-parameter MoE with a separate 196B Engram conditional-memory pool, activating 16B per token while generating (8B during prefill), with a 1M-token context window and native image input through its DeepSeek-ViT encoder.
Compressed Sparse Attention 2 and SWA Bounded Replay cut the global KV cache to 890 bytes per token — roughly a quarter of DeepSeek V4 Flash — so whole-repo prompts and long agent sessions stay fast and affordable at full 1M context.
The result is a clear step up in agentic coding at Flash cost: 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1 at maximum reasoning effort, up from 82.7 and 54.4 for V4 Flash, with function calling, JSON output, and a 1–100 reasoning-effort dial — including none to skip thinking — to pay for depth only where a request needs it.
1Mtoken context window
Sparse attention keeps the global KV cache at 890 bytes per token — about a quarter of DeepSeek V4 Flash.
90.6on Terminal-Bench 2.1
Up from 82.7 for DeepSeek V4 Flash at maximum reasoning effort — from MIT-licensed open weights.
96%avg. cache hit rate on Entrim
Measured on coding-agent workloads — repeated context bills at the cache-read rate.
DeepSeek V4.1 Flash is strongest where agents reason over a lot of context and now look at screens too: coding loops, tool chains, and document pipelines that need long memory at Flash cost.
Holds a plan across long tool-call chains — 74.2 on DeepSWE 1.1, up from 54.4 for V4 Flash, and a 3471 Codeforces rating.
A from-scratch DeepSeek-ViT encoder reads screenshots, diagrams, and scanned pages directly — 95.6 on DocVQA for the base model — with no separate vision model in the loop.
An integer from 1 to 100 (with low, high, and max presets) sets thinking depth per request, and effort none skips reasoning entirely — pay for depth only where a request needs it.
Strict function calling and JSON mode at every reasoning effort — tool arguments and typed pipelines that parse the first time.
Long-running coding agents
Plans, edits, and verifies across long tool-call chains — 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE 1.1 at maximum reasoning effort.
Agents grounded in screenshots and diagrams
Native image input reads UI screenshots, charts, and scanned pages straight from the prompt, with no separate vision model in the loop.
Whole-repo and long-document analysis
The 1M-token window holds a full repository or document set in one prompt, and sparse attention keeps full-context requests affordable.
Browser and workflow automation
54.8 on AutomationBench — up from 37.7 for V4 Flash — for multi-step tool use across real applications.
Fact recall without retrieval
Standalone knowledge recall is still limited for a sparse-activation model (42.3 on SimpleQA-Verified for the base model) — pair it with RAG for factual workloads.
Frontier-grade open-ended reasoning
Open-ended exam-style reasoning is not the strength — 36.8 on Humanity’s Last Exam without tools. Give research-style questions retrieval and tools, where it reaches 63.9 on HLE.
Audio or video inputs
Inputs are text and images only. Use GLM 5.3 Flash for video, and transcribe audio upstream before sending it to the model.
median 131.3tokens/s
median 1,247ms
90.6Terminal-Bench 2.1
Agentic coding and shell tasks run end-to-end in a real terminal.
74.2DeepSWE 1.1
Resolves real software-engineering issues across long tool-call chains.
54.8AutomationBench
Tool-driven automation across real applications and workflows.
Code examples
curl -X POST "https://api.entrim.ai/v1/chat/completions" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $ENTRIM_API_KEY" \ -d '{ "model": "deepseek-ai/DeepSeek-V4.1-Flash", "messages": [ { "role": "user", "content": "Why do they call it a building if it is already built?" } ] }'SDKs & docs
Use familiar SDK patterns with quickstart examples for common production setups.
OpenAI-compatible
Keep the request format your team already knows.
No cold starts
Production requests are served without model spin-up delays.
High-throughput agentic work with 1M context
Multimodal coding agents with 1M context
Dense multimodal coding with thinking control