23 models, one API key
Access DeepSeek V4, Nemotron 3, Llama, Mistral, GPT-OSS, FLUX and more frontier models through one OpenAI-compatible API key — every model priced per token, no subscriptions. Every model below is health-checked continuously, so the catalog only lists what is actually running. Filter by capability, copy the model id, and start calling Speka in minutes.
What models can I call on Speka?
Speka is a unified AI model API that provides developer access to 23 frontier open-weight models — including DeepSeek V4, Nemotron 3 Ultra 550B, Mistral Nemotron, GPT-OSS and FLUX.1 for image generation — through a single OpenAI-compatible endpoint. Pricing is transparent and per-token, with no subscriptions or seat fees. Input prices run from $0.01 per million tokens for embeddings to $0.90 per million tokens for Nemotron 3 Ultra 550B. Developers switch between models by changing one parameter — the model field — with no SDK changes required. Model weights and benchmarks for these open models are published by their labs — for example DeepSeek and Black Forest Labs (FLUX) on Hugging Face.
GLM 5.2
Z.ai's flagship bilingual model with strong agentic tool-use and long-form writing in both English and Chinese.
z-ai/glm-5.2Llama 3.1 8B Instruct
Fast, cheap and capable. Ideal for high-volume classification, routing and lightweight chat.
meta/llama-3.1-8b-instructLlama 3.2 3B Instruct
The smallest Llama instruct model — the cheapest way to run classification, routing and short replies at scale.
meta/llama-3.2-3b-instructLlama 3.3 70B Instruct
Meta's flagship 70B instruct model — 405B-class quality at a fraction of the cost. A dependable default for production chat.
meta/llama-3.3-70b-instructMiniMax M3
MiniMax's large mixture-of-experts assistant — strong long-context comprehension and agentic workflows.
minimaxai/minimax-m3Mistral Nemotron
Mistral and NVIDIA's joint instruct model: excellent function calling, 80+ languages and consistently low latency for production chat.
mistralai/mistral-nemotronNemotron 3 Ultra 550B
NVIDIA's flagship open model: 550B parameters with 55B active per token. Frontier quality for demanding generation, analysis and agentic work.
nvidia/nemotron-3-ultra-550b-a55bStep 3.7 Flash
Latency-optimised assistant model that thinks briefly before answering — a good fit for interactive products.
stepfun-ai/step-3.7-flashGPT-OSS 20B
The small GPT-OSS tier — quick code completion, refactors and shell/tool calls at a fraction of the price.
openai/gpt-oss-20bLaguna XS 2.1
Poolside's compact software-engineering model, tuned for repository-aware code completion and edits with very low latency.
poolside/laguna-xs-2.1Llama Nemotron Embed 1B v2
Current-generation NVIDIA retrieval embedding model — 2048-dimension vectors tuned for multilingual RAG.
nvidia/llama-nemotron-embed-1b-v2Nemotron 3 Embed 1B
Nemotron 3 embedding model for semantic search, clustering and reranking-style similarity scoring.
nvidia/nemotron-3-embed-1bNV-Embed v1
High-accuracy general-purpose text embeddings for semantic search and clustering.
nvidia/nv-embed-v1NV-EmbedQA E5 v5
Robust retrieval embeddings tuned for question answering and RAG pipelines.
nvidia/nv-embedqa-e5-v5FLUX.1 [dev]
High-fidelity text-to-image generation with excellent prompt adherence and typography.
black-forest-labs/flux.1-devInkling
Auto-discovered from the NVIDIA catalog on 2026-07-28.
thinkingmachines/inklingNemotron 3 Nano 30B
Compact MoE reasoning model (30B total / 3B active). Sub-second answers on maths, logic and structured extraction at a fraction of the cost.
nvidia/nemotron-3-nano-30b-a3bNemotron 3 Super 120B
NVIDIA's mixture-of-experts reasoning model — 120B total parameters with 12B active, so frontier-class reasoning arrives at small-model latency.
nvidia/nemotron-3-super-120b-a12bNemotron Nano 9B
Compact reasoning model that punches above its weight on math and coding benchmarks. Fast and cheap.
nvidia/nvidia-nemotron-nano-9b-v2Nemotron Super 49B
NVIDIA's reasoning-tuned Nemotron model — strong math, logic and agentic tool-use at an efficient size.
nvidia/llama-3.3-nemotron-super-49b-v1.5Llama 3.2 11B Vision
Lightweight vision-language model for fast image captioning, OCR and visual Q&A.
meta/llama-3.2-11b-vision-instructLlama 3.2 90B Vision
Large multimodal model for image understanding, document Q&A, charts and visual reasoning.
meta/llama-3.2-90b-vision-instructNemotron Nano 12B VL
NVIDIA's compact vision-language model — fast document, chart and screenshot understanding for automation pipelines.
nvidia/nemotron-nano-12b-v2-vl