Serverless API Pricing

Pay less.
Run faster.

Usage-based pricing on every open-source model, with sub-second responses and predictable latency under production load.

All prices in USD per 1M tokens.

ModelInput Cached Output
$0.10$0.0175$0.35
$0.20$0.005$0.60
$0.10$0.04$0.40
$0.10$0.04$0.30
$0.02——
$0.029$0.01$0.13
$0.09$0.015$0.17

Why teams run inference on Entrim

Lower cost per request without giving up speed, privacy, or the integration you already have.

Speed from the first token

Low latency and high output speed on GPUs we own and tune, never on capacity resold by a middleman.

Lower cost by engineering

Lower request cost from an inference runtime we engineered for efficiency, not from cheaper models.

Privacy that passes review

A European company running its own datacenter. No prompt or output logging, no training on your data.

Built for production traffic

OpenAI-compatible, so it plugs right into the SDKs, agents, and harnesses your team already works with.

Up to 80% Lower Cost

Same models.
Same tokens.
Lower bill.

Compare Entrim’s pricing, powered by an optimized inference runtime, against other providers using the same token counts per request.

Estimate your Savings

Select a model

▾

Token usage9B input - 909M output

10B
ProviderInput / 1MOutput / 1MMonthly est.
AkashML$0.45$3.20
$7,000.00
io.net$0.44$3.15
$6,863.64
Chutes$0.40$3.00
$6,363.64
Entrim$0.10$0.40
$1,272.73 Save up to 80%

Need more than serverless?

Higher rate limits, a clear path to dedicated capacity when you outgrow serverless, and support for DPAs, Zero Data Retention, and security reviews.

FAQ

Here are the most common questions users ask before getting started.

Run your workload on Entrim

Bring your prompts, token counts, and latency targets. You will know if it fits before you move traffic.
All services are online

© 2026. Entrim. All Rights Reserved.

Privacy policy•Terms of service