Compare the best 80 vision LLMs and try 75 of them on your own images, free in the Roboflow Playground. 31 are open-weight, so you can self-host them for free under their licenses.
80 models · 31 open-weight · 75 free to try · prices synced Sep 12, 2026
Benchmark Leader
GPT-6 Astra
Highest score on our vision benchmark across six tasks, 53 vision language models tested
Open-Weight Leader
Qwen3.8 27B
Best score among the open-weight models, #17 of 53 overall on our vision benchmark across six tasks
Fastest on Benchmark
Gemini 3.5 Flash-Lite
Quickest in our vision benchmark across six tasks runs: 2.7s per request on average
Lowest Measured Cost
Qwen3.7 Flash
Cheapest in our vision benchmark across six tasks runs: $0.0001 per request as measured, not list price
Ranked on our vision benchmark across six tasks, 53 vision language models tested, updated Sep 5, 2026
31 models with downloadable weights you can self-host under their licenses (MIT, Apache 2.0, and Modified MIT). 26 run live in the Playground through hosted APIs, so self-hosting is optional.
| Actions | ||||
|---|---|---|---|---|
GLM 5.3 FlashZ.ai | 320B total, 18B active | MIT | Aug 2026 | |
Qwen3.8 27BQwen | 27.78B | Apache 2.0 | Aug 2026 | |
Muse Glimmer 30BMeta | 29.6B | Apache 2.0 | Aug 2026 | |
Kimi K3Moonshot AI | 2.8T | Modified MIT | Jul 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
Qwen3.5-27BQwen | 27B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Llama 4 MaverickMeta | 400B | Custom | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Custom | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Custom | Mar 2025 | |
Gemma 3 27BGoogle | — | Custom | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Custom | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
LLaVA-1.5Microsoft | 7B, 13B | Custom | Oct 2023 | |
Qwen-VLQwen | — | Custom | Aug 2023 |
49 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
| Actions | |||||
|---|---|---|---|---|---|
GPT-6 AstraOpenAI NEW | $10.00 | $50.00 | 1.1M | Sep 2026 | |
Gemini 3.8 FlashGoogle NEW | $0.75 | $3.75 | 1.0M | Sep 2026 | |
Muse Spark 1.3Meta NEW | $1.25 | $4.25 | 1.0M | Sep 2026 | |
Claude Fable 5.1Anthropic NEW | $10.00 | $50.00 | 1M | Sep 2026 | |
Qwen3.8 FlashQwen | $0.15 | $0.47 | 1M | Aug 2026 | |
Gemini 3.7 FlashGoogle | $0.75 | $3.75 | 1.0M | Aug 2026 | |
Grok 4.6SpaceXAI | $2.00 | $6.00 | 500K | Aug 2026 | |
Muse Spark 1.2Meta | $1.25 | $4.25 | 1.0M | Aug 2026 | |
Qwen3.8 MaxQwen | — | — | 984K | Aug 2026 | |
Qwen3.7 FlashQwen | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle | $0.75 | $3.75 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI | $0.20 | $1.20 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI | $2.00 | $10.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI | $2.00 | $12.00 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Grok 4.5SpaceXAI | $2.00 | $6.00 | 500K | Jul 2026 | |
Claude Sonnet 5Anthropic | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GLM 5V TurboZ.ai | $1.20 | $4.00 | 200K | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4SpaceXAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
The expensive mistake in either direction: using a frontier vision LLM for a job a small trained model does better and a thousand times cheaper, or spending weeks building a training pipeline for a job a VLM handles with one prompt. The decision comes down to how open-ended the task is.
If the task is fixed and repetitive (detect these products, classify these defects, count people in this camera feed), a trained detector or classifier wins on every production axis: per-image cost measured in fractions of a cent, millisecond latency, deployment on edge hardware, and consistency you can put error bars on. Vision LLMs answer the same question in seconds for orders of magnitude more money, with answers that can vary between calls.
Open-ended instructions, language output, and world knowledge are the VLM-only zone: describe this scene for an incident report, extract these fields from any invoice layout, answer whatever question a user types about an image. No trained classifier can do these, because the label space is unbounded. If the question changes per request, or the output is prose or structured text, you need the language model.
Within the family, capability tracks reasoning depth more than vision. Reading a receipt or writing alt text works on small open models at a fraction of frontier pricing; multi-step document reasoning and unusual images separate the top models quickly. Our vision benchmark on the evals page scores the models we have tested across six task types; the verdict module above names the current leader.
Mature production systems use the VLM as the reasoning layer, not the whole pipeline: a trained detector localizes at volume, OCR reads at volume, and the VLM handles the open-ended residue (judgment calls, summaries, escalations). Prototype everything with the VLM first, then move the fixed, high-volume steps to specialized models as the requirements stop changing.
The bottom line: Fixed question at volume: train a classic model. Open-ended question or language output: vision LLM. At scale, the answer is usually a pipeline with both.
LLMs with vision capabilities (also called vision language models or VLMs) are large language models extended with an image encoder, so they accept images as input alongside text. The encoder converts the image into token embeddings the language model reads the same way it reads words, which lets one model describe scenes, answer questions about images, read and transcribe text, and return structured output like JSON from a visual input. They inherit the strengths and weaknesses of their language side: broad world knowledge and instruction following, but also a tendency to state wrong visual details fluently, and looser spatial precision than dedicated detectors. This page lists 80 vision LLMs, including 31 open-weight options you can self-host; 75 of them run live in the Playground so you can test them on your own images.
Every vision LLM is multimodal, but not every multimodal model is an LLM. Multimodal is the broader property of handling more than one input type. CLIP is multimodal because it aligns images and text in one embedding space, yet it generates no language at all. A vision LLM is specifically a language model that can see: multimodal input, conversational text output.
On our vision benchmark across six tasks (53 models tested, updated Sep 5, 2026), GPT-6 Astra by OpenAI currently scores highest at 86.6%. The full ranking is on our evals page. For fixed categories in production, a model fine-tuned on your own data still often wins.
Yes. 31 of the 80 models here are open-weight (for example GLM 5.3 Flash, Qwen3.8 27B, and Muse Glimmer 30B), free to self-host under their licenses (MIT, Apache 2.0, and Modified MIT).
Yes. You can run 75 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 80 vision LLMs in the Roboflow Playground catalog: 31 open-weight models you can self-host and 49 proprietary models accessed through provider APIs; 75 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.