Model Library

Open-source models for production inference

Compare available models by family, workload fit, pricing, context length, and performance profile - then test them through Entrim's OpenAI-compatible API.

Explore all models

10 open-source models across 6 families, served through one OpenAI-compatible API.

DeepSeek V4 Flash

Version 0731

High-throughput agentic work with 1M context

1M
FP4
Thinking
Tool Calling
Input / M
$0.09
Output / M
$0.17
Cached / M
$0.015

Tool-heavy coding at 3B-active cost

262K
FP8
Vision
Multimodal
Video
Input / M
$0.05
Output / M
$0.25
Cached / M
$0.025

Dense agentic coding with thinking control

262K
FP8
Vision
Multimodal
Video
Input / M
$0.10
Output / M
$0.40
Cached / M
$0.04

Gemma 4 quality at 4B-active speed

256K
FP8
Vision
Multilingual
Input / M
$0.05
Output / M
$0.25
Cached / M
$0.025

Multimodal long-context work with thinking control

256K
FP8
Vision
Multilingual
Input / M
$0.10
Output / M
$0.30
Cached / M
$0.04

Top-ranked multilingual embeddings for search

32K
Embeddings
Multilingual
Input / M
$0.02
Output / M
Cached / M

o4-mini-class reasoning for tool-driven coding

131.1K
FP4
Thinking
Input / M
$0.029
Output / M
$0.13
Cached / M
$0.01

Marathon coding agents with visual context

262K
FP4
Vision
Multimodal
Input / M
$0.6
Output / M
$3.2
Cached / M
$0.10

Low-latency reasoning on a tiny budget

131.1K
FP4
Thinking
Input / M
$0.02
Output / M
$0.10
Cached / M
$0.005

Long-horizon engineering with planning depth

200K
FP8
Thinking
Input / M
$0.9
Output / M
$3.0
Cached / M
$0.15

Test the model that fits your workload.

Use $25 free credits to compare models with your prompts, token sizes, latency targets, and cost assumptions before moving traffic.

All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policyTerms of service