visionMeta

Llama 3.2 11B Vision

Lightweight vision-language model for fast image captioning, OCR and visual Q&A.

meta/llama-3.2-11b-vision-instruct
Running· verified 3 min ago · 174ms

What is Llama 3.2 11B Vision?

Llama 3.2 11B Vision is Meta's vision model. Lightweight vision-language model for fast image captioning, OCR and visual Q&A. It runs on Speka's OpenAI-compatible /v1/chat/completions endpoint at $0.06 / 1M input and $0.06 / 1M output tokens, using the model ID meta/llama-3.2-11b-vision-instruct.

Test it live

Live playgroundLlama 3.2 11B Visionin-browser demo

Test Llama 3.2 11B Vision live

Send a message and watch the model respond. Multi-turn — it remembers the conversation.