Explore vision AI models from every major lab and try them on your own images: object detection, OCR, classification, captioning, segmentation, and more. Every model is free to try in the Playground, many are free and open-source to deploy, and new models are added as they ship.
124 models · 25 tasks · latest model added July 27, 2026
84 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 34 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle NEW | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
SAM 3Meta | — | Custom | Nov 2025 | |
SAM 3D ObjectsMeta | — | Custom | Nov 2025 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
RF-DETR SegmentationRoboflow | 33.6M-38.6M | Apache 2.0 | Oct 2025 | |
YOLO26Ultralytics | 2.4M-55.7M | AGPL 3.0 | Oct 2025 | |
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
RF-DETRRoboflow | 30.5M-126.9M | Apache 2.0 | Mar 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
YOLOETHU-MIG | 10M-50M | AGPL 3.0 | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
YOLOv12THU-MIG | 2.6M-59.1M | AGPL 3.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
DEIMIntellindust AI Lab | 4M-62M | Apache 2.0 | Dec 2024 | |
D-FINEUSTC | 4M-62M | Apache 2.0 | Oct 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
YOLO11Ultralytics | 2.6M-56.9M | AGPL 3.0 | Sep 2024 | |
| 38.9M-224.4M | Apache 2.0 | Jul 2024 | ||
Depth Anything V2ByteDance | 25M to 1.3B | Apache 2.0 | Jun 2024 | |
YOLOv10THU-MIG | 2.3M-29.5M | AGPL 3.0 | May 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
YOLO WorldTencent AI Lab | 13M | GPL v3 | Feb 2024 | |
YOLOv9Academia Sinica | 2.0M-57.3M | GPL v3 | Feb 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
Grounded SAMIDEA Research | — | Apache 2.0 | Jan 2024 | |
SuryaMindee | — | GPL v3 | Jan 2024 | |
SAM-CLIPApple | — | Custom | Oct 2023 | |
LLaVA-1.5Microsoft | 7B, 13B | Custom | Oct 2023 | |
Qwen-VLQwen | — | Custom | Aug 2023 | |
YOLO-NASDeci AI | — | Custom | May 2023 | |
RT-DETRBaidu | 20M-76M | Apache 2.0 | Apr 2023 | |
DINOv2Meta | 21M-1.1B | Apache 2.0 | Apr 2023 | |
YOLOv8 Pose EstimationUltralytics | 2.9M-57.6M | AGPL 3.0 | Apr 2023 | |
| 91M-636M | Apache 2.0 | Apr 2023 | ||
SigLIPGoogle | 200M-900M | Apache 2.0 | Mar 2023 | |
Grounding DINOIDEA Research | 172M-341M | Apache 2.0 | Mar 2023 | |
YOLOv8Ultralytics | 3.2M-68.2M | AGPL 3.0 | Jan 2023 | |
YOLOv8 ClassificationUltralytics | — | AGPL 3.0 | Jan 2023 | |
YOLOv8 Instance SegmentationUltralytics | 2.7M-62.8M | AGPL 3.0 | Jan 2023 | |
RTMDetOpenMMLab | 4.8M-94.9M | GPL v3 | Dec 2022 | |
Co-DETROpenMMLab | 304M | MIT | Nov 2022 | |
YOLOv7Academia Sinica | 6.2M-151.7M | GPL v3 | Jul 2022 | |
OWL-ViTGoogle | — | Apache 2.0 | May 2022 | |
ByteTrackByteDance | — | MIT | Oct 2021 | |
TrOCRMicrosoft | 61.4M-600M | MIT | Sep 2021 | |
YOLOXMegvii | 0.91M-99.1M | Apache 2.0 | Jul 2021 | |
YOLOSHugging Face | — | MIT | Jun 2021 | |
CLIPOpenAI | — | MIT | Feb 2021 | |
docTRMindee | — | Apache 2.0 | Feb 2021 | |
YOLOv4-tinyAcademia Sinica | — | Custom | Nov 2020 | |
Vision Transformer (ViT)Google | 86M-632M | Apache 2.0 | Oct 2020 | |
DETRMeta | ~41M | Apache 2.0 | May 2020 | |
YOLOv4Academia Sinica | — | — | Apr 2020 | |
YOLOv5Ultralytics | 1.9M-86.7M | AGPL 3.0 | Jan 2020 | |
EfficientDetGoogle | 3.9M-51.9M | Apache 2.0 | Nov 2019 | |
Detectron2Meta | — | Apache 2.0 | Sep 2019 | |
MediaPipeGoogle | — | Apache 2.0 | Jul 2019 | |
MobileNet SSD v2Google | 15.3M | MIT | Jan 2018 | |
MobileNetV2Google | ~3.4M | Apache 2.0 | Jan 2018 | |
Mask R-CNNMeta | 44.4M | MIT | Oct 2017 | |
ResNet-32Meta | 0.46M | MIT | Dec 2015 | |
ResNet-34Meta | 21.8M | MIT | Dec 2015 | |
ResNet-50Microsoft | 25.6M | MIT | Dec 2015 | |
Faster R-CNNMicrosoft | 41.8M | MIT | Jun 2015 |
40 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 | |
Google Vision OCRGoogle | — | — | — | Feb 2016 |
Vision AI models are models that understand visual input: images, video, and documents. They span two families. Vision language models (GPT, Claude, Gemini, Qwen VL) reason about images with natural language and handle open-ended tasks zero-shot. Specialized computer vision models (RF-DETR, SAM 3, CLIP, Florence-2) are built for one task and run faster, cheaper, and often more accurately on it in production.
Prototype with a VLM to validate the task, then move to a specialized or fine-tuned model when classes are fixed and volume grows. For object detection, that usually means training RF-DETR on your own data. In a Roboflow Workflow you can chain both: a detector localizes, a VLM interprets only the crops that matter.
Vision AI models are models that interpret visual data such as images, video, and documents. They include vision language models that answer questions about images in natural language, and specialized computer vision models for tasks like object detection, OCR, classification, and segmentation.
All major frontier models accept images: GPT-5.6, Claude Opus 4.7, and Gemini 3.1 Pro, plus open-weight options like Qwen3 VL, Gemma 4, and Pixtral. This page lists 100+ vision-capable models you can test on your own image right now.
Yes. Many models on this page are free and open-source, including CLIP and Florence-2 (MIT), RF-DETR and Qwen3 VL (Apache 2.0). They run with no API costs when self-hosted via Roboflow Inference. Every model here is also free to try in the Playground.
Computer vision models are the specialized family within vision AI: models built and trained for a specific visual task. Vision AI also includes multimodal language models that handle visual tasks through prompting. If you want to train a model on your own data, see the Roboflow model library; this page is for trying hosted models.
This page lists 100+ vision AI models from Google, OpenAI, Anthropic, Meta, Qwen, Mistral, and more, covering object detection, OCR, classification, captioning, and segmentation. Every model can be tested on your own image in the browser and compared head to head before you commit to one.