Catalog

23 models, one API key

Access DeepSeek V4, Nemotron 3, Llama, Mistral, GPT-OSS, FLUX and more frontier models through one OpenAI-compatible API key — every model priced per token, no subscriptions. Every model below is health-checked continuously, so the catalog only lists what is actually running. Filter by capability, copy the model id, and start calling Speka in minutes.

What models can I call on Speka?

Speka is a unified AI model API that provides developer access to 23 frontier open-weight models — including DeepSeek V4, Nemotron 3 Ultra 550B, Mistral Nemotron, GPT-OSS and FLUX.1 for image generation — through a single OpenAI-compatible endpoint. Pricing is transparent and per-token, with no subscriptions or seat fees. Input prices run from $0.01 per million tokens for embeddings to $0.90 per million tokens for Nemotron 3 Ultra 550B. Developers switch between models by changing one parameter — the model field — with no SDK changes required. Model weights and benchmarks for these open models are published by their labs — for example DeepSeek and Black Forest Labs (FLUX) on Hugging Face.

Z

GLM 5.2

Z.ai
chat

Z.ai's flagship bilingual model with strong agentic tool-use and long-form writing in both English and Chinese.

z-ai/glm-5.2
Input / 1M
$0.60
Output / 1M
$2.20
Open model
Meta logo

Llama 3.1 8B Instruct

Meta
chat

Fast, cheap and capable. Ideal for high-volume classification, routing and lightweight chat.

meta/llama-3.1-8b-instruct
Input / 1M
$0.05
Output / 1M
$0.05
Open model
Meta logo

Llama 3.2 3B Instruct

Meta
chat

The smallest Llama instruct model — the cheapest way to run classification, routing and short replies at scale.

meta/llama-3.2-3b-instruct
Input / 1M
$0.03
Output / 1M
$0.03
Open model
Meta logo

Llama 3.3 70B Instruct

Meta
chat

Meta's flagship 70B instruct model — 405B-class quality at a fraction of the cost. A dependable default for production chat.

meta/llama-3.3-70b-instruct
Input / 1M
$0.20
Output / 1M
$0.20
Open model
M

MiniMax M3

MiniMax
chat

MiniMax's large mixture-of-experts assistant — strong long-context comprehension and agentic workflows.

minimaxai/minimax-m3
Input / 1M
$0.30
Output / 1M
$1.20
Open model
Mistral AI logo

Mistral Nemotron

Mistral AI
chat

Mistral and NVIDIA's joint instruct model: excellent function calling, 80+ languages and consistently low latency for production chat.

mistralai/mistral-nemotron
Input / 1M
$0.30
Output / 1M
$0.30
Open model
NVIDIA logo

Nemotron 3 Ultra 550B

NVIDIA
chat

NVIDIA's flagship open model: 550B parameters with 55B active per token. Frontier quality for demanding generation, analysis and agentic work.

nvidia/nemotron-3-ultra-550b-a55b
Input / 1M
$0.90
Output / 1M
$2.70
Open model
S

Step 3.7 Flash

StepFun
chat

Latency-optimised assistant model that thinks briefly before answering — a good fit for interactive products.

stepfun-ai/step-3.7-flash
Input / 1M
$0.25
Output / 1M
$0.75
Open model
O

GPT-OSS 20B

OpenAI
code

The small GPT-OSS tier — quick code completion, refactors and shell/tool calls at a fraction of the price.

openai/gpt-oss-20b
Input / 1M
$0.07
Output / 1M
$0.30
Open model
P

Laguna XS 2.1

Poolside
code

Poolside's compact software-engineering model, tuned for repository-aware code completion and edits with very low latency.

poolside/laguna-xs-2.1
Input / 1M
$0.20
Output / 1M
$0.60
Open model
NVIDIA logo

Llama Nemotron Embed 1B v2

NVIDIA
embedding

Current-generation NVIDIA retrieval embedding model — 2048-dimension vectors tuned for multilingual RAG.

nvidia/llama-nemotron-embed-1b-v2
Input / 1M
$0.012
Open model
NVIDIA logo

Nemotron 3 Embed 1B

NVIDIA
embedding

Nemotron 3 embedding model for semantic search, clustering and reranking-style similarity scoring.

nvidia/nemotron-3-embed-1b
Input / 1M
$0.012
Open model
NVIDIA logo

NV-Embed v1

NVIDIA
embedding

High-accuracy general-purpose text embeddings for semantic search and clustering.

nvidia/nv-embed-v1
Input / 1M
$0.016
Open model
NVIDIA logo

NV-EmbedQA E5 v5

NVIDIA
embedding

Robust retrieval embeddings tuned for question answering and RAG pipelines.

nvidia/nv-embedqa-e5-v5
Input / 1M
$0.01
Open model
Black Forest Labs logo

FLUX.1 [dev]

Black Forest Labs
image

High-fidelity text-to-image generation with excellent prompt adherence and typography.

black-forest-labs/flux.1-dev
Per image
$0.04
Open model
TM

Inkling

Thinking Machines
reasoning

Auto-discovered from the NVIDIA catalog on 2026-07-28.

thinkingmachines/inkling
Input / 1M
$0.20
Output / 1M
$0.20
Open model
NVIDIA logo

Nemotron 3 Nano 30B

NVIDIA
reasoning

Compact MoE reasoning model (30B total / 3B active). Sub-second answers on maths, logic and structured extraction at a fraction of the cost.

nvidia/nemotron-3-nano-30b-a3b
Input / 1M
$0.15
Output / 1M
$0.20
Open model
NVIDIA logo

Nemotron 3 Super 120B

NVIDIA
reasoning

NVIDIA's mixture-of-experts reasoning model — 120B total parameters with 12B active, so frontier-class reasoning arrives at small-model latency.

nvidia/nemotron-3-super-120b-a12b
Input / 1M
$0.40
Output / 1M
$0.60
Open model
NVIDIA logo

Nemotron Nano 9B

NVIDIA
reasoning

Compact reasoning model that punches above its weight on math and coding benchmarks. Fast and cheap.

nvidia/nvidia-nemotron-nano-9b-v2
Input / 1M
$0.10
Output / 1M
$0.10
Open model
NVIDIA logo

Nemotron Super 49B

NVIDIA
reasoning

NVIDIA's reasoning-tuned Nemotron model — strong math, logic and agentic tool-use at an efficient size.

nvidia/llama-3.3-nemotron-super-49b-v1.5
Input / 1M
$0.35
Output / 1M
$0.40
Open model
Meta logo

Llama 3.2 11B Vision

Meta
vision

Lightweight vision-language model for fast image captioning, OCR and visual Q&A.

meta/llama-3.2-11b-vision-instruct
Input / 1M
$0.06
Output / 1M
$0.06
Open model
Meta logo

Llama 3.2 90B Vision

Meta
vision

Large multimodal model for image understanding, document Q&A, charts and visual reasoning.

meta/llama-3.2-90b-vision-instruct
Input / 1M
$0.35
Output / 1M
$0.40
Open model
NVIDIA logo

Nemotron Nano 12B VL

NVIDIA
vision

NVIDIA's compact vision-language model — fast document, chart and screenshot understanding for automation pipelines.

nvidia/nemotron-nano-12b-v2-vl
Input / 1M
$0.10
Output / 1M
$0.15
Open model
FAQ

Frequently asked questions

Speka provides 23 frontier open-weight models across six categories: reasoning (DeepSeek V4 Flash, DeepSeek V4 Pro, Nemotron 3 Super 120B, Nemotron 3 Nano 30B, Nemotron Super 49B, Nemotron Nano 9B), chat (Nemotron 3 Ultra 550B, Llama 3.3 70B, Llama 3.1 8B, Llama 3.2 3B, Mistral Nemotron, Mistral Medium 3.5, GLM 5.2, MiniMax M3, Step 3.7 Flash), code (GPT-OSS 120B, GPT-OSS 20B, Laguna XS 2.1), vision (Llama 3.2 90B Vision, Nemotron Nano 12B VL, Llama 3.2 11B Vision), embeddings (NV-EmbedQA E5 v5, Llama Nemotron Embed 1B v2, Nemotron 3 Embed 1B, NV-Embed v1), and image generation (FLUX.1 [dev], FLUX.1 [schnell]). The live catalog is verified continuously — see the model list below for what is running right now.
DeepSeek V4 Flash on Speka costs $0.27 per million input tokens and $1.10 per million output tokens. Pricing is per-token with no subscription or seat fee — you pay only for what you use, with no minimum commitment.
Yes. Speka uses an OpenAI-compatible endpoint, so any application already using the OpenAI Python SDK, TypeScript SDK, or HTTP API can switch to Speka by changing two values: the base URL and the model ID. No other code changes are required.
GPT-OSS 120B is the flagship code model in the catalog — a 120-billion-parameter open-weight model strong at code generation, refactoring, and tool-use, priced at $0.15 per million input tokens and $0.60 per million output tokens. GPT-OSS 20B ($0.07 / $0.30 per million tokens) and Poolside Laguna XS 2.1 cover latency-sensitive completion, and Llama 3.1 8B Instruct at $0.05 per million tokens is the cheapest capable alternative.
Yes. Speka includes four embedding models: NV-EmbedQA E5 v5, tuned for question-answering retrieval at $0.01 per million tokens; Llama Nemotron Embed 1B v2 and Nemotron 3 Embed 1B, current-generation multilingual retrieval models at $0.012 per million tokens; and NV-Embed v1, a general-purpose model at $0.016 per million tokens. All return standard embedding vectors compatible with any vector database. For the generation step, any Speka chat or reasoning model is available under the same API key.