Compare the best 67 vision LLMs and try 62 of them on your own image, free in the Roboflow Playground. 28 are open-weight, so you can self-host them for free under their licenses.
67 models · 28 open-weight · 62 free to try · prices synced Jul 29, 2026
Benchmark Leader
Gemini 3.5 Flash
Highest score on our vision benchmark across six tasks, 23 models tested
Open-Weight Leader
Kimi K3
Best score among the open-weight models, #15 of 23 overall on our vision benchmark across six tasks
Fastest on Benchmark
Gemini 3.5 Flash-Lite
Quickest in our vision benchmark across six tasks runs: 2.7s per request on average
Lowest Measured Cost
Qwen 3.7 Flash
Cheapest in our vision benchmark across six tasks runs: $0.0001 per request as measured, not list price
Ranked on our vision benchmark across six tasks, 23 models tested, updated Jul 28, 2026
28 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 23 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
LLaVA-1.5Microsoft | 7B, 13B | Custom | Oct 2023 | |
Qwen-VLQwen | — | Custom | Aug 2023 |
39 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
The expensive mistake in either direction: using a frontier vision LLM for a job a small trained model does better and a thousand times cheaper, or spending weeks building a training pipeline for a job a VLM handles with one prompt. The decision comes down to how open-ended the task is.
If the task is fixed and repetitive (detect these products, classify these defects, count people in this camera feed), a trained detector or classifier wins on every production axis: per-image cost measured in fractions of a cent, millisecond latency, deployment on edge hardware, and consistency you can put error bars on. Vision LLMs answer the same question in seconds for orders of magnitude more money, with answers that can vary between calls.
Open-ended instructions, language output, and world knowledge are the VLM-only zone: describe this scene for an incident report, extract these fields from any invoice layout, answer whatever question a user types about an image. No trained classifier can do these, because the label space is unbounded. If the question changes per request, or the output is prose or structured text, you need the language model.
Within the family, capability tracks reasoning depth more than vision. Reading a receipt or writing alt text works on small open models at a fraction of frontier pricing; multi-step document reasoning and unusual images separate the top models quickly. Our vision benchmark on the evals page scores the models we have tested across six task types; the verdict module above names the current leader.
Mature production systems use the VLM as the reasoning layer, not the whole pipeline: a trained detector localizes at volume, OCR reads at volume, and the VLM handles the open-ended residue (judgment calls, summaries, escalations). Prototype everything with the VLM first, then move the fixed, high-volume steps to specialized models as the requirements stop changing.
The bottom line: Fixed question at volume: train a classic model. Open-ended question or language output: vision LLM. At scale, the answer is usually a pipeline with both.
LLMs with vision capabilities (also called vision language models or VLMs) are large language models extended with an image encoder, so they accept images as input alongside text. The encoder converts the image into token embeddings the language model reads the same way it reads words, which lets one model describe scenes, answer questions about images, read and transcribe text, and return structured output like JSON from a visual input. They inherit the strengths and weaknesses of their language side: broad world knowledge and instruction following, but also a tendency to state wrong visual details fluently, and looser spatial precision than dedicated detectors. This page lists 67 vision LLMs, including 28 open-weight options you can self-host; 62 of them run live in the Playground so you can test them on your own images.
Every vision LLM is multimodal, but not every multimodal model is an LLM. Multimodal is the broader property of handling more than one input type. CLIP is multimodal because it aligns images and text in one embedding space, yet it generates no language at all. A vision LLM is specifically a language model that can see: multimodal input, conversational text output.
On our vision benchmark across six tasks (23 models tested, updated Jul 28, 2026), Gemini 3.5 Flash by Google currently scores highest at 86.6%. The full ranking is on our evals page. For fixed categories in production, a model fine-tuned on your own data still often wins.
Yes. 28 of the 67 models here are open-weight (for example Kimi K3, Qwen3.6 27B, and Qwen3.6 35B A3B), free to self-host under their licenses (Modified MIT, Apache 2.0, and MIT).
Yes. You can run 62 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 67 vision LLMs in the Roboflow Playground catalog: 28 open-weight models you can self-host and 39 proprietary models accessed through provider APIs; 62 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.