Pricing
Pay per token. Start free.
Free tier
$0
- ✓ $1 of inference a day, any model
- ✓ 60 requests a minute
- ✓ No card, no forms
- ✓ US or Canada
Pay as you go
From $0.09 / 1M tokens
Per model, per token
Input, cached input and output, priced for each model. Full list below.
No router fee
model: "auto" bills at the price of whichever model served the request.
Same in both regions
US and Canada cost the same. No residency premium.
Hard caps, no surprises
Your human approves one monthly cap. Keys can carry their own budgets.
Every model
What each model costs.
USD per 1M tokens · cheapest first
| Model | Context | Input | Cached input | Output |
|---|---|---|---|---|
| Gemma 4 31BGoogle DeepMind | 256K | $0.09 | $0.05 | $0.34 |
| GLM-5.3-FlashZ.ai | 1M | $0.15 | $0.03 | $0.50 |
| K2 Horizon 375BIFM | 512K | $0.50 | — | $0.50 |
| Mistral Small 4Mistral AI | 256K | $0.15 | — | $0.60 |
| gpt-oss-120bOpenAI | 128K | $0.15 | $0.015 | $0.60 |
| DeepSeek-V4.1-FlashDeepSeek | 1M | $0.22 | $0.007 | $0.66 |
| Llama 4 MaverickMeta | 1M | $0.27 | — | $0.85 |
| MiMo-V2.6-ProXiaomi | 1M | $0.43 | — | $0.87 |
| OLMo 3 32BAi2 | — | $0.90 | — | $0.90 |
| MiniMax-M3MiniMax | 1M | $0.30 | $0.06 | $1.20 |
| Muse GlimmerMeta | — | $0.35 | $0.04 | $1.50 |
| Mistral Large 3Mistral AI | 256K | $0.50 | — | $1.50 |
| Qwen3.8-27BAlibaba | 256K | $0.15 | — | $1.88 |
| Nemotron 3 UltraNVIDIA | — | $0.60 | $0.12 | $2.40 |
| Hy4 previewTencent | 1M | $0.834 | $0.042 | $2.50 |
| DeepSeek V4 ProDeepSeek | 1M | $1.32 | $0.044 | $3.96 |
| InklingThinking Machines | 1M | $1.00 | $0.17 | $4.05 |
| GLM-5.3Z.ai | 1M | $1.40 | $0.26 | $4.40 |
| Qwen3.8-2.4TAlibaba | 961K | $2.00 | $0.25 | $6.00 |
| Kimi K3Moonshot AI | 1M | $3.00 | $0.30 | $15.00 |
Fine-tuning
Train once. Pay base price after.
LoRA adapters, billed per 1M training tokens (dataset tokens × epochs). Your fine-tuned model is private to your account and costs the same per token as its base model. Download the adapter any time.
| Base model | Training / 1M tokens |
|---|---|
| Qwen3.8-27BAlibaba | $0.50 |
| Gemma 4 31BGoogle DeepMind | $0.50 |
| OLMo 3 32BAi2 | $0.50 |
| Muse GlimmerMeta | $0.50 |
| gpt-oss-120bOpenAI | $3.00 |
| Mistral Small 4Mistral AI | $3.00 |
Rate limits
Published, not hidden.
Free
60
requests / min
200,000 tokens / min
$1/day per account, no card
Pay as you go
3,000
requests / min
5,000,000 tokens / min
Monthly cap you approve
Scale
Custom
requests / min
Custom tokens / min
Talk to us
Questions
Billing, plainly.
- Do I need a card to start?
- No. The free tier gives every account $1 of inference a day on any model, at the normal per-token prices, with no card on file. Your agent can set it up in one command.
- What needs a paid plan?
- Batch inference, fine-tuning, higher rate limits and more than five keys. Your human approves a monthly cap once; nothing is ever charged beyond it.
- How does billing work past the free tier?
- Your agent runs billing upgrade, which returns one approval link for your human. They add a card and a monthly cap. You're billed monthly for what you used, never more than the cap.
- What does the router cost?
- Nothing extra. With model: "auto" you pay the per-token price of whichever model served each request.
- What are cached input tokens?
- Repeated prompt prefixes (system prompts, long documents) are cached automatically and billed at the model's lower cached-input rate.
- How are prices set?
- Per model, benchmarked against comparable serverless prices for the same open weights. When a lab ships a new model, it gets its own price on day one.
- Does US or Canada cost more?
- No. Same price in both regions.
- Do you sell dedicated GPUs?
- No. Everything is serverless and per-token, so there's nothing to reserve or size.