Compare the best 75 foundation vision models and try 63 of them on your own image, free in the Roboflow Playground. 36 are open-weight, so you can self-host them for free under their licenses.
75 models · 36 open-weight · 63 free to try · prices synced Jul 28, 2026
36 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and Custom). 24 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
SAM 3Meta | — | Custom | Nov 2025 | |
SAM 3D ObjectsMeta | — | Custom | Nov 2025 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
Depth Anything V2ByteDance | 25M to 1.3B | Apache 2.0 | Jun 2024 | |
SAM-CLIPApple | — | Custom | Oct 2023 | |
DINOv2Meta | 21M-1.1B | Apache 2.0 | Apr 2023 | |
| 91M-636M | Apache 2.0 | Apr 2023 | ||
SigLIPGoogle | 200M-900M | Apache 2.0 | Mar 2023 | |
Grounding DINOIDEA Research | 172M-341M | Apache 2.0 | Mar 2023 | |
OWL-ViTGoogle | — | Apache 2.0 | May 2022 | |
CLIPOpenAI | — | MIT | Feb 2021 | |
DETRMeta | ~41M | Apache 2.0 | May 2020 | |
Detectron2Meta | — | Apache 2.0 | Sep 2019 | |
Mask R-CNNMeta | 44.4M | MIT | Oct 2017 |
39 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
Foundation models differ most in what they output, and that, not accuracy, is the first decision: language, embeddings, or masks. Pick the wrong output type and no amount of model quality saves the integration.
Need descriptions, answers, or structured text? That is the vision LLM subset. Need vectors for search, similarity, or clustering? Contrastive encoders (CLIP, SigLIP) when queries include text, self-supervised encoders (DINOv2) when precision on pure visuals matters. Need pixel-accurate regions? The SAM family. These subfamilies are not interchangeable: an embedding model cannot explain an image, and a VLM makes an expensive, imprecise segmenter.
Foundation models earn the name by adapting without full retraining, but the paths differ in cost and control. Prompting is instant and flexible (VLMs, SAM-family). Frozen features plus a small trained head is the cheap, stable middle path for classification and retrieval on encoder models. Fine-tuning buys domain accuracy at the price of training infrastructure and maintenance. Start with the cheapest path that could work; escalate only when its ceiling shows.
Foundation models frequently end up as components: SAM labels the masks that train your compact production segmenter, a VLM pre-labels the images that train your detector, CLIP embeddings bootstrap your first classifier. When you evaluate one, weigh it as a data engine and prototyping tool, not only as the thing that serves production traffic; the compact model it helps you train is often what actually ships.
The bottom line: Choose the output type first (language, vectors, or masks), then the cheapest adaptation path that could work. The foundation model that ships is often the one that trained your production model.
Foundation vision models are large models pretrained on web-scale image or image-text data that serve as general-purpose bases rather than single-task tools. The class spans vision language models (GPT and Gemini class systems, open VLMs like Qwen VL), contrastive encoders like CLIP and SigLIP, self-supervised encoders like DINOv2, and promptable segmenters like the SAM family. What unites them is transfer: they carry enough general visual knowledge to be applied to new problems through prompting, fine-tuning, or as frozen feature extractors, instead of being trained from scratch per task. This page lists 75 foundation vision models, including 36 open-weight options you can self-host; 63 of them run live in the Playground so you can test them on your own images.
No. Language-capable VLMs are the most visible members, but CLIP, DINOv2, and SAM are foundation vision models with no language decoder: they embed, match, or segment without generating a word of text. If you need conversational answers about an image, you want the vision LLM subset; if you need embeddings, similarity, or masks at scale, the non-LLM foundation models are usually smaller, faster, and cheaper.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the foundation vision models on this page and try them on your own image to see which fits.
Yes. 36 of the 75 models here are open-weight (for example Kimi K3, Qwen3.6 27B, and Qwen3.6 35B A3B), free to self-host under their licenses (Modified MIT, Apache 2.0, and Custom).
Yes. You can run 63 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 75 foundation vision models in the Roboflow Playground catalog: 36 open-weight models you can self-host and 39 proprietary models accessed through provider APIs; 63 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.