Compare the best 75 multimodal vision models and try 65 of them on your own image, free in the Roboflow Playground. 36 are open-weight, so you can self-host them for free under their licenses.
75 models · 36 open-weight · 65 free to try · prices synced Jul 28, 2026
36 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 26 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle NEW | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
SAM 3Meta | — | Custom | Nov 2025 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
YOLO WorldTencent AI Lab | 13M | GPL v3 | Feb 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
Grounded SAMIDEA Research | — | Apache 2.0 | Jan 2024 | |
LLaVA-1.5Microsoft | 7B, 13B | Custom | Oct 2023 | |
Qwen-VLQwen | — | Custom | Aug 2023 | |
SigLIPGoogle | 200M-900M | Apache 2.0 | Mar 2023 | |
CLIPOpenAI | — | MIT | Feb 2021 |
39 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
Multimodal models come in two working styles, and most integration pain comes from picking the wrong one: generative models that read inputs and write answers, and contrastive models that score how well an image and a text belong together.
Vision LLMs are the right tool when the output is consumed as language: descriptions, answers, extracted fields, reports. They are slower and priced per token, and their failure mode is a fluent wrong answer, so pair them with checks where correctness matters. Everything on this page that chats, captions, or answers questions sits in this group.
CLIP-style models never write a sentence; they place images and text in one vector space. That single trick powers zero-shot classification, text-to-image search, deduplication, and content routing at a per-image cost and latency generative models cannot approach. If the output feeds an index, a threshold, or a ranking function rather than a person, the contrastive family is usually the right and radically cheaper answer.
A common production shape: a contrastive model narrows millions of images to candidates (search, filtering, routing), and a generative model reasons over the shortlist. Sizing each stage to its job keeps quality where users see it and cost where volume lives.
The bottom line: Output read by a person: generative. Output consumed by a system: contrastive. High-volume products usually chain the cheap one into the smart one.
Multimodal vision models process more than one kind of input, most commonly images together with text, and in newer systems video, audio, or documents in the same context. Architecturally that means separate encoders per modality feeding a shared representation, either a generative language model (vision LLMs) or a joint embedding space (contrastive models like CLIP). The practical payoff is reasoning across modalities in one call: answer a text question about an image, ground a phrase in a photo, or search images with language. This page lists 75 multimodal vision models, including 36 open-weight options you can self-host; 65 of them run live in the Playground so you can test them on your own images.
No. Generative multimodal models (vision LLMs) converse and describe, but contrastive multimodal models like CLIP and SigLIP only score how well an image and a text match, which is exactly what zero-shot classification and cross-modal search need. The chat interface is one way to be multimodal, not the definition of it.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the multimodal vision models on this page and try them on your own image to see which fits.
Yes. 36 of the 75 models here are open-weight (for example Kimi K3, Gemma 4 12B, and Qwen3.6 27B), free to self-host under their licenses (Modified MIT, Apache 2.0, and MIT).
Yes. You can run 65 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 75 multimodal vision models in the Roboflow Playground catalog: 36 open-weight models you can self-host and 39 proprietary models accessed through provider APIs; 65 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.