Pricing

One key, every model, no overage penalties.

Every plan includes a monthly usage allowance measured in real model spend. Exceed it and you only pay standard per-token rates — no overage penalties, no surprises.

What does Speka's AI API cost?

Speka is a unified AI API gateway that gives developers access to 27 frontier models — including DeepSeek V4, Nemotron 3 Ultra 550B, Mistral Nemotron, GPT-OSS 120B, and Llama — through a single API key and one billing account. Pricing is usage-based: each plan includes a monthly token allowance at real model rates, and usage beyond that allowance bills at standard per-token rates with no overage penalties and no automatic upgrades. The Free plan costs $0/month with $1 of included usage and no credit card required; paid plans start at $19/month.

Free

$0.00 /mo

Kick the tires. No card required.

  • $1 of model usage / month
  • 10 requests / minute
  • 1 API key
  • Access to all open models
  • Web playground
  • Community support

Starter

$19.00 /mo

For indie hackers and side projects.

  • $25 of model usage included
  • 60 requests / minute
  • 5 API keys
  • All models incl. vision & image
  • Usage analytics & alerts
  • Email support
  • Pay-as-you-go overage

Scale

$399.00 /mo

For high-volume, latency-sensitive workloads.

  • $750 of model usage included
  • 1,200 requests / minute
  • 200 API keys
  • Highest-priority routing
  • Volume discounts on overage
  • Dedicated support channel
  • SSO & audit logs (on request)
  • Custom SLAs available

Prices in USD. Cancel or change plans any time. Need higher volume or an SLA? Talk to us.

Per-token model rates

Usage is billed against your plan's included allowance at these rates (USD per 1M tokens).

ModelPublisherInput / 1MOutput / 1M
DeepSeek V4 Flash
deepseek-ai/deepseek-v4-flash
DeepSeek logoDeepSeek$0.27$1.10
DeepSeek V4 Pro
deepseek-ai/deepseek-v4-pro
DeepSeek logoDeepSeek$0.55$2.20
Nemotron 3 Super 120B
nvidia/nemotron-3-super-120b-a12b
NVIDIA logoNVIDIA$0.40$0.60
Nemotron 3 Nano 30B
nvidia/nemotron-3-nano-30b-a3b
NVIDIA logoNVIDIA$0.15$0.20
Nemotron Super 49B
nvidia/llama-3.3-nemotron-super-49b-v1.5
NVIDIA logoNVIDIA$0.35$0.40
Nemotron Nano 9B
nvidia/nvidia-nemotron-nano-9b-v2
NVIDIA logoNVIDIA$0.10$0.10
Nemotron 3 Ultra 550B
nvidia/nemotron-3-ultra-550b-a55b
NVIDIA logoNVIDIA$0.90$2.70
Llama 3.3 70B Instruct
meta/llama-3.3-70b-instruct
Meta logoMeta$0.20$0.20
Llama 3.1 8B Instruct
meta/llama-3.1-8b-instruct
Meta logoMeta$0.05$0.05
Llama 3.2 3B Instruct
meta/llama-3.2-3b-instruct
Meta logoMeta$0.03$0.03
Mistral Nemotron
mistralai/mistral-nemotron
Mistral AI logoMistral AI$0.30$0.30
Mistral Medium 3.5
mistralai/mistral-medium-3.5-128b
Mistral AI logoMistral AI$0.40$0.40
GLM 5.2
z-ai/glm-5.2
ZZ.ai$0.60$2.20
MiniMax M3
minimaxai/minimax-m3
MMiniMax$0.30$1.20
Step 3.7 Flash
stepfun-ai/step-3.7-flash
SStepFun$0.25$0.75
GPT-OSS 120B
openai/gpt-oss-120b
OOpenAI$0.15$0.60
GPT-OSS 20B
openai/gpt-oss-20b
OOpenAI$0.07$0.30
Laguna XS 2.1
poolside/laguna-xs-2.1
PPoolside$0.20$0.60
Llama 3.2 90B Vision
meta/llama-3.2-90b-vision-instruct
Meta logoMeta$0.35$0.40
Nemotron Nano 12B VL
nvidia/nemotron-nano-12b-v2-vl
NVIDIA logoNVIDIA$0.10$0.15
Llama 3.2 11B Vision
meta/llama-3.2-11b-vision-instruct
Meta logoMeta$0.06$0.06

Embeddings are billed per 1M input tokens and image models per image — see the full catalog.

FAQ

Frequently asked questions

Speka offers four usage-based plans: Free ($0/month, includes $1 of model usage, no credit card required), Starter ($19/month), Pro ($99/month, 99.9% uptime target), and Scale ($399/month, volume discounts on overage). All plans include access to the full model catalog. Usage beyond the included allowance bills at per-token rates with no overage penalties.
Yes. The Free plan costs $0/month and includes $1 of monthly model usage with no credit card required. It provides access to all open models in the catalog, 10 requests per minute, 1 API key, and a web playground — designed for evaluation, prototyping, and personal projects.
DeepSeek V4 Flash (deepseek-ai/deepseek-v4-flash) on Speka is priced at $0.27 per 1M input tokens and $1.10 per 1M output tokens. Usage is first applied against your plan's monthly allowance; any usage beyond it bills at the same per-token rate with no penalty or tier change.
When you exceed your plan's monthly token allowance, additional usage is billed at standard per-token rates for each model. There are no overage penalties, no automatic plan upgrades, and no minimum overage charges. Scale plan customers receive volume discounts on overage usage. You pay exactly what you use — nothing more.
Yes. Plans can be changed or cancelled at any time — no lock-in, no cancellation fees. Prices are in USD. Usage is billed against your plan's included monthly allowance at standard per-token rates.

Start free and let your usage decide

Start on the free plan — no card required — and upgrade only when you need more.