Roboflow

Vision AI Models

Explore vision AI models from every major lab and try them on your own images: object detection, OCR, classification, captioning, segmentation, and more. The catalog spans general-purpose vision language models and specialized computer vision models built for a single task. Every model is free to try in the Playground, many are free and open-source to deploy, and new models are added as they ship.

Best Vision Language Models 2026

The best vision language model on our vision benchmark across six tasks right now is GPT-6 Astra by OpenAI, scoring 86.6% across 53 vision language models tested, followed by Gemini 3.5 Flash at 86.0%. Updated Sep 5, 2026. See the full ranking.

Benchmark scores

Ranked on our vision benchmark across six tasks, 53 vision language models tested, updated Sep 5, 2026

#1OpenAIGPT-6 AstraAPI86.6%
#2GoogleGemini 3.5 FlashAPI86.0%
#3GoogleGemini 3.7 FlashAPI85.2%
#4GoogleGemini 3.8 FlashAPI85.1%
#5QwenQwen3.8 MaxAPI83.9%

See the full benchmark

Explore Vision AI Models

138 models · 25 tasks · latest model added September 3, 2026

Modality:Multimodal
Actions
OpenAINEW
ProprietarySep 2026
Try
GoogleNEW
ProprietarySep 2026
Try
MetaNEW
ProprietarySep 2026
Try
AnthropicNEW
ProprietarySep 2026
Try
MITAug 2026
Try
CustomAug 2026
Try
Qwen
Apache 2.0Aug 2026
Try
Google
ProprietaryAug 2026
Try
Grok
Grok 4.6SpaceXAI
ProprietaryAug 2026
Try
Apache 2.0Aug 2026
Try
ProprietaryAug 2026
Try
Qwen
Apache 2.0Aug 2026
Try
ProprietaryJul 2026
Try
Anthropic
Claude Opus 5Anthropic
ProprietaryJul 2026
Try
ProprietaryJul 2026
Try
Google
ProprietaryJul 2026
Try
MoonshotAI
Kimi K3Moonshot AI
Modified MITJul 2026
Try
OpenAI
ProprietaryJul 2026
Try
OpenAI
ProprietaryJul 2026
Try
OpenAI
ProprietaryJul 2026
Try
ProprietaryJul 2026
Try
Grok
Grok 4.5SpaceXAI
ProprietaryJul 2026
Try
Anthropic
ProprietaryJun 2026
Try
Anthropic
ProprietaryJun 2026
Try
Google
Apache 2.0Jun 2026
Anthropic
ProprietaryMay 2026
Try
Google
ProprietaryMay 2026
Try
OpenAI
GPT-5.5OpenAI
ProprietaryApr 2026
Try
Qwen
Apache 2.0Apr 2026
Anthropic
ProprietaryApr 2026
Try
Apache 2.0Apr 2026
Try
ProprietaryApr 2026
Google
Apache 2.0Apr 2026
Try
Google
Apache 2.0Apr 2026
Try
Qwen
ProprietaryApr 2026
Z.ai
ProprietaryApr 2026
Try
OpenAI
ProprietaryMar 2026
Try
OpenAI
ProprietaryMar 2026
Try
Z.ai
MITMar 2026
Try
OpenAI
GPT-5.4OpenAI
ProprietaryMar 2026
Try
ProprietaryMar 2026
Try
Qwen
Apache 2.0Mar 2026
Apache 2.0Feb 2026
Apache 2.0Feb 2026
Qwen
Apache 2.0Feb 2026
Try
Google
ProprietaryFeb 2026
Try
Anthropic
ProprietaryFeb 2026
Try
Apache 2.0Feb 2026
Anthropic
ProprietaryFeb 2026
Try
MoonshotAI
Kimi K2.5Moonshot AI
Modified MITJan 2026
Google
ProprietaryDec 2025
Try
OpenAI
GPT-5.2OpenAI
ProprietaryDec 2025
Try
Anthropic
ProprietaryNov 2025
Try
Meta
SAM 3Meta
CustomNov 2025
Try
OpenAI
GPT-5.1OpenAI
ProprietaryNov 2025
Try
Anthropic
ProprietaryOct 2025
Try
Apache 2.0Oct 2025
Apache 2.0Oct 2025
Anthropic
ProprietarySep 2025
Try
Apache 2.0Sep 2025
Try
Mistral
ProprietaryAug 2025
OpenAI
GPT-5OpenAI
ProprietaryAug 2025
Try
OpenAI
ProprietaryAug 2025
Try
OpenAI
ProprietaryAug 2025
Try
ProprietaryJul 2025
Try
Google
ProprietaryJul 2025
Try
Grok
Grok 4SpaceXAI
ProprietaryJul 2025
Azure
Florence-2Microsoft
MITJun 2025
Try
Google
ProprietaryJun 2025
Try
CustomApr 2025
CustomApr 2025
Mistral
Apache 2.0Mar 2025
Google
CustomMar 2025
Google
CustomMar 2025
Google
CustomMar 2025
HuggingFace
SmolVLM2Hugging Face
Apache 2.0Feb 2025
Qwen
ProprietaryFeb 2025
Apache 2.0Jan 2025
Google
CustomDec 2024
Mistral
Apache 2.0Sep 2024
Google
PaliGemmaGoogle
CustomMay 2024
Tencent
YOLO WorldTencent AI Lab
GPL v3Feb 2024
Try
Moondream
Moondream 2Moondream
Apache 2.0Jan 2024
IDEA Research
Grounded SAMIDEA Research
Apache 2.0Jan 2024
Azure
LLaVA-1.5Microsoft
CustomOct 2023
Qwen
CustomAug 2023
Google
SigLIPGoogle
Apache 2.0Mar 2023
OpenAI
CLIPOpenAI
MITFeb 2021

What Are Vision AI Models?

Vision AI models are models that understand visual input: images, video, and documents. They span two families. Vision language models (GPT, Claude, Gemini, Qwen VL) reason about images with natural language and handle open-ended tasks zero-shot. Specialized computer vision models (RF-DETR, SAM 3, CLIP, Florence-2) are built for one task and run faster, cheaper, and often more accurately on it in production.

Which family do you need?

Prototype with a VLM to validate the task, then move to a specialized or fine-tuned model when classes are fixed and volume grows. For object detection, that usually means training RF-DETR on your own data. In a Roboflow Workflow you can chain both: a detector localizes, a VLM interprets only the crops that matter.

Frequently Asked Questions About Vision AI Models

Vision AI models are models that interpret visual data such as images, video, and documents. They include vision language models that answer questions about images in natural language, and specialized computer vision models for tasks like object detection, OCR, classification, and segmentation.

All major frontier models accept images: GPT-5.6, Claude Opus 4.7, and Gemini 3.1 Pro, plus open-weight options like Qwen3 VL, Gemma 4, and Pixtral. This page lists 100+ vision-capable models you can test on your own image right now.

Yes. Many models on this page are free and open-source, including CLIP and Florence-2 (MIT), RF-DETR and Qwen3 VL (Apache 2.0). They run with no API costs when self-hosted via Roboflow Inference. Every model here is also free to try in the Playground.

Computer vision models are the specialized family within vision AI: models built and trained for a specific visual task. Vision AI also includes multimodal language models that handle visual tasks through prompting. If you want to train a model on your own data, see the Roboflow model library; this page is for trying hosted models.

This page lists 100+ vision AI models from Google, OpenAI, Anthropic, Meta, Qwen, Mistral, and more, covering object detection, OCR, classification, captioning, and segmentation. Every model can be tested on your own image in the browser and compared head to head before you commit to one.