10 models, one API key
Access DeepSeek V4, Nemotron 3, Llama, Mistral, GPT-OSS, FLUX and more frontier models through one OpenAI-compatible API key — every model priced per token, no subscriptions. Every model below is health-checked continuously, so the catalog only lists what is actually running. Filter by capability, copy the model id, and start calling Speka in minutes.
What models can I call on Speka?
Speka is a unified AI model API that provides developer access to 10 frontier open-weight models — including DeepSeek V4, Nemotron 3 Ultra 550B, Mistral Nemotron, GPT-OSS and FLUX.1 for image generation — through a single OpenAI-compatible endpoint. Pricing is transparent and per-token, with no subscriptions or seat fees. Input prices run from $0.01 per million tokens for embeddings to $0.90 per million tokens for Nemotron 3 Ultra 550B. Developers switch between models by changing one parameter — the model field — with no SDK changes required. Model weights and benchmarks for these open models are published by their labs — for example DeepSeek and Black Forest Labs (FLUX) on Hugging Face.
Nemotron 3 Ultra 550B
NVIDIA's flagship open model: 550B parameters with 55B active per token. Frontier quality for demanding generation, analysis and agentic work.
nvidia/nemotron-3-ultra-550b-a55bStep 3.7 Flash
Latency-optimised assistant model that thinks briefly before answering — a good fit for interactive products.
stepfun-ai/step-3.7-flashGPT-OSS 120B
Open-weight 120B model with strong code generation, refactoring and tool-use. Built for coding agents and automation.
openai/gpt-oss-120bGPT-OSS 20B
The small GPT-OSS tier — quick code completion, refactors and shell/tool calls at a fraction of the price.
openai/gpt-oss-20bNemotron 3 Embed 1B
Nemotron 3 embedding model for semantic search, clustering and reranking-style similarity scoring.
nvidia/nemotron-3-embed-1bFLUX.1 [dev]
High-fidelity text-to-image generation with excellent prompt adherence and typography.
black-forest-labs/flux.1-devNemotron 3 Nano 30B
Compact MoE reasoning model (30B total / 3B active). Sub-second answers on maths, logic and structured extraction at a fraction of the cost.
nvidia/nemotron-3-nano-30b-a3bNemotron 3 Super 120B
NVIDIA's mixture-of-experts reasoning model — 120B total parameters with 12B active, so frontier-class reasoning arrives at small-model latency.
nvidia/nemotron-3-super-120b-a12bLlama 3.2 11B Vision
Lightweight vision-language model for fast image captioning, OCR and visual Q&A.
meta/llama-3.2-11b-vision-instructLlama 3.2 90B Vision
Large multimodal model for image understanding, document Q&A, charts and visual reasoning.
meta/llama-3.2-90b-vision-instruct