ScaledThought

Pricing

Pay per token. Start free.

Every model has its own per-token price, the same in the US and Canada. No router fee, no minimums, and no capacity to buy.

Free tier

$0

  • ✓ $1 of inference a day, any model
  • ✓ 60 requests a minute
  • ✓ No card, no forms
  • ✓ US or Canada

Pay as you go

From $0.09 / 1M tokens

Per model, per token

Input, cached input and output, priced for each model. Full list below.

No router fee

model: "auto" bills at the price of whichever model served the request.

Same in both regions

US and Canada cost the same. No residency premium.

Hard caps, no surprises

Your human approves one monthly cap. Keys can carry their own budgets.

Every model

What each model costs.

USD per 1M tokens · cheapest first

ModelContextInputCached inputOutput
Gemma 4 31BGoogle DeepMind256K$0.09$0.05$0.34
GLM-5.3-FlashZ.ai1M$0.15$0.03$0.50
K2 Horizon 375BIFM512K$0.50—$0.50
Mistral Small 4Mistral AI256K$0.15—$0.60
gpt-oss-120bOpenAI128K$0.15$0.015$0.60
DeepSeek-V4.1-FlashDeepSeek1M$0.22$0.007$0.66
Llama 4 MaverickMeta1M$0.27—$0.85
MiMo-V2.6-ProXiaomi1M$0.43—$0.87
OLMo 3 32BAi2—$0.90—$0.90
MiniMax-M3MiniMax1M$0.30$0.06$1.20
Muse GlimmerMeta—$0.35$0.04$1.50
Mistral Large 3Mistral AI256K$0.50—$1.50
Qwen3.8-27BAlibaba256K$0.15—$1.88
Nemotron 3 UltraNVIDIA—$0.60$0.12$2.40
Hy4 previewTencent1M$0.834$0.042$2.50
DeepSeek V4 ProDeepSeek1M$1.32$0.044$3.96
InklingThinking Machines1M$1.00$0.17$4.05
GLM-5.3Z.ai1M$1.40$0.26$4.40
Qwen3.8-2.4TAlibaba961K$2.00$0.25$6.00
Kimi K3Moonshot AI1M$3.00$0.30$15.00

Fine-tuning

Train once. Pay base price after.

LoRA adapters, billed per 1M training tokens (dataset tokens × epochs). Your fine-tuned model is private to your account and costs the same per token as its base model. Download the adapter any time.

Base modelTraining / 1M tokens
Qwen3.8-27BAlibaba$0.50
Gemma 4 31BGoogle DeepMind$0.50
OLMo 3 32BAi2$0.50
Muse GlimmerMeta$0.50
gpt-oss-120bOpenAI$3.00
Mistral Small 4Mistral AI$3.00

Rate limits

Published, not hidden.

Free

60

requests / min

200,000 tokens / min

$1/day per account, no card

Pay as you go

3,000

requests / min

5,000,000 tokens / min

Monthly cap you approve

Scale

Custom

requests / min

Custom tokens / min

Talk to us

Questions

Billing, plainly.

Do I need a card to start?
No. The free tier gives every account $1 of inference a day on any model, at the normal per-token prices, with no card on file. Your agent can set it up in one command.
What needs a paid plan?
Batch inference, fine-tuning, higher rate limits and more than five keys. Your human approves a monthly cap once; nothing is ever charged beyond it.
How does billing work past the free tier?
Your agent runs billing upgrade, which returns one approval link for your human. They add a card and a monthly cap. You're billed monthly for what you used, never more than the cap.
What does the router cost?
Nothing extra. With model: "auto" you pay the per-token price of whichever model served each request.
What are cached input tokens?
Repeated prompt prefixes (system prompts, long documents) are cached automatically and billed at the model's lower cached-input rate.
How are prices set?
Per model, benchmarked against comparable serverless prices for the same open weights. When a lab ships a new model, it gets its own price on day one.
Does US or Canada cost more?
No. Same price in both regions.
Do you sell dedicated GPUs?
No. Everything is serverless and per-token, so there's nothing to reserve or size.

Opens your agent with the prompt typed. You press Enter.