All prices in USD per 1M tokens.
| Model | InputInput price / M | CachedCache input price / M | OutputOutput price / M | |
|---|---|---|---|---|
GLM 5.3 Flash | $0.10 | $0.0175 | $0.35 | Details |
DeepSeek V4.1 Flash | $0.20 | $0.005 | $0.60 | Details |
Qwen 3.8 27B | $0.10 | $0.04 | $0.40 | Details |
Gemma 4 31B | $0.10 | $0.04 | $0.30 | Details |
Qwen3 Embedding 8B | $0.02 | — | — | Details |
GPT OSS 120B | $0.029 | $0.01 | $0.13 | Details |
DeepSeek V4 Flash | $0.09 | $0.015 | $0.17 | Details |
Low latency and high output speed on GPUs we own and tune, never on capacity resold by a middleman.
Lower request cost from an inference runtime we engineered for efficiency, not from cheaper models.
A European company running its own datacenter. No prompt or output logging, no training on your data.
OpenAI-compatible, so it plugs right into the SDKs, agents, and harnesses your team already works with.
Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.
Estimate your Savings
Select a model
Token usage(Usage split, 10:1: 9B input - 909M output)
| Provider | Input / 1M | Output / 1M | Monthly est. |
|---|---|---|---|
| AkashML | $0.45 | $3.20 | $7,000.00 |
| io.net | $0.44 | $3.15 | $6,863.64 |
| Chutes | $0.40 | $3.00 | $6,363.64 |
| Entrim | $0.10 | $0.40 | $1,272.73 Save up to 80% |