Compare the best 69 visual question answering models and try 62 of them on your own image, free in the Roboflow Playground. 30 are open-weight, so you can self-host them for free under their licenses.
69 models · 30 open-weight · 62 free to try · prices synced Jul 28, 2026
The best visual question answering model on our visual reasoning benchmark right now is Gemini 3.5 Flash by Google, scoring 84.1% across 23 models tested, followed by Gemini 3.6 Flash at 80.1%. Updated Jul 28, 2026. See the full ranking.
Benchmark Leader
Gemini 3.5 Flash
Highest score on our visual reasoning benchmark, 23 models tested
Open-Weight Leader
Kimi K3
Best score among the open-weight models, #16 of 23 overall on our visual reasoning benchmark
Fastest on Benchmark
Qwen3-VL 235B
Quickest in our visual reasoning benchmark runs: 1.7s per request on average
Lowest Measured Cost
Qwen 3.7 Flash
Cheapest in our visual reasoning benchmark runs: <$0.0001 per request as measured, not list price
Ranked on our visual reasoning benchmark, 23 models tested, updated Jul 28, 2026
30 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 23 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle NEW | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
LLaVA-1.5Microsoft | 7B, 13B | Custom | Oct 2023 | |
Qwen-VLQwen | — | Custom | Aug 2023 |
39 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
VQA quality tracks general reasoning ability more than any other vision task: the model must read the question, find the evidence in the image, and reason to an answer. That makes the choice mostly about how hard your questions are and how expensive a wrong answer is.
Questions that chain steps (counting plus comparison, reading text then interpreting it, spatial reasoning in cluttered scenes) separate the top models from the rest quickly. If your questions are open-ended or your users type anything they want, the frontier tier is the safe default; our visual reasoning benchmark on the evals page shows how the current models actually rank on this.
When the question is fixed and simple ("is the shelf empty?", "is a person present?"), a small open VLM answers at a fraction of the cost, and often a trained classifier or detector answers it even more reliably: a fixed question is really a classification or detection task in disguise, and converting it usually improves both accuracy and unit economics.
VQA models answer fluently whether or not they are right. For production use, prefer questions whose answers can be checked (counts you can verify with a detector, text you can verify with OCR), ask for the evidence in the answer, and treat free-form answers about safety-critical content as drafts for review rather than decisions.
The bottom line: Open-ended questions need a frontier model; a fixed, repeated question is usually a classification or detection task wearing a question mark, and converting it is cheaper and more reliable.
Visual question answering (VQA) is the task of answering a free-form question about an image, which forces a model to combine perception with language understanding and reasoning. A single question may require recognizing objects, reading text inside the image, counting, comparing positions, or chaining several of those steps, and the output is text rather than a fixed label or box. Modern vision language models handle VQA zero-shot: the image and question go in together and the model generates the answer. Benchmarks score exact-match or judged answer quality across question types. VQA drives document and screenshot understanding, assistive technology for low-vision users, and ad hoc extraction like "what is the total on this receipt?". This page lists 69 visual question answering models, including 30 open-weight options you can self-host; 62 of them run live in the Playground so you can test them on your own images.
On our visual reasoning benchmark (23 models tested, updated Jul 28, 2026), Gemini 3.5 Flash by Google currently scores highest at 84.1%. The full ranking is on our evals page. For fixed categories in production, a model fine-tuned on your own data still often wins.
Yes. 30 of the 69 visual question answering models here are open-weight (for example Kimi K3, Gemma 4 12B, and Qwen3.6 27B), free to self-host under their licenses (Modified MIT, Apache 2.0, and MIT).
Yes. You can run 62 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 69 visual question answering models in the Roboflow Playground catalog: 30 open-weight models you can self-host and 39 proprietary models accessed through provider APIs; 62 of them run live in the Roboflow Playground on your own images. On our visual reasoning benchmark (23 models tested), Gemini 3.5 Flash by Google currently scores highest. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.