embeddingNVIDIA

Llama Nemotron Embed 1B v2

Current-generation NVIDIA retrieval embedding model — 2048-dimension vectors tuned for multilingual RAG.

nvidia/llama-nemotron-embed-1b-v2
Running· verified just now · 133ms

What is Llama Nemotron Embed 1B v2?

Llama Nemotron Embed 1B v2 is NVIDIA's embedding model. Current-generation NVIDIA retrieval embedding model — 2048-dimension vectors tuned for multilingual RAG. It runs on Speka's OpenAI-compatible /v1/embeddings endpoint at $0.012 per 1M input tokens, using the model ID nvidia/llama-nemotron-embed-1b-v2.

Test it live

Live playgroundLlama Nemotron Embed 1B v2in-browser demo

Embed text with Llama Nemotron Embed 1B v2

Enter text to see a vector preview.