Compare all 25 Google vision models we track, 15 of them open-weight. Try 15 of them on your own images, free in the Roboflow Playground.
25 models · 15 open-weight · 15 free to try · Updated Jul 2026
Google ships vision models on two tiers, and our catalog tracks 25 of them. 15 are open-weight: the weights are downloadable, so you can self-host under their own licenses. The other 10 are proprietary, reached through an API and billed by the provider. The open-weight side runs on Apache 2.0, Custom, and MIT licenses. The newest additions are Gemini 3.5 Flash-Lite (Jul 2026) and Gemini 3.6 Flash (Jul 2026). The lineup reaches back to Google Vision OCR, released Feb 2016.
Between them, the Google models we list cover 17 distinct vision tasks. The widest coverage is vision language (18 models), captioning (17), and classification (17). For vision language, start with Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, or Gemma 4 12B.
On the open-weight tier, Gemma 4 31B is the largest at 31B parameters and MobileNetV2 the smallest at ~3.4M, small enough to run on a single GPU or on-device. 10 of the 15 open Google models carry a permissive license, so commercial use is straightforward; the other 5 ship under terms worth reading before you deploy.
On the API tier, Google's most recent entry is Gemini 3.5 Flash-Lite (Jul 2026), billed per token by the provider rather than run on your own hardware. Which tier to start on is a constraints question, not a quality one: reach for the open-weight side when you need offline inference, predictable per-image cost, or a checkpoint you can fine-tune on your own data, and for the API side when you want broad general reasoning without managing GPUs. 15 of the 25 Google vision models run live in the Roboflow Playground, so you can run the same image through several of them and compare the answers before committing to one.
What each of the 25 Google vision models in our catalog is built for, and how you run it.
15 models with downloadable weights you can self-host under their licenses (Apache 2.0, Custom, and MIT). 5 run live in the Playground through hosted APIs, so self-hosting is optional.
Gemma 4 12BGoogle | 12B | Apache 2.0 | Jun 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Gemma 3 12BGoogle | 12B | Custom | Mar 2025 | |
Gemma 3 27BGoogle | — | Custom | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Custom | Mar 2025 | |
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
SigLIPGoogle | 200M-900M | Apache 2.0 | Mar 2023 | |
OWL-ViTGoogle | — | Apache 2.0 | May 2022 | |
Vision Transformer (ViT)Google | 86M-632M | Apache 2.0 | Oct 2020 | |
EfficientDetGoogle | 3.9M-51.9M | Apache 2.0 | Nov 2019 | |
MediaPipeGoogle | — | Apache 2.0 | Jul 2019 | |
MobileNet SSD v2Google | 15.3M | MIT | Jan 2018 | |
MobileNetV2Google | ~3.4M | Apache 2.0 | Jan 2018 |
10 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle | $1.50 | $7.50 | 1M | Jul 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Google Vision OCRGoogle | — | — | — | Feb 2016 |
18 of the 25 Google vision models we track handle vision language: Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemma 4 12B, Gemini 3.5 Flash, Gemma 4 26B A4B, and Gemma 4 31B, and 12 more. Each model page lists its full task coverage, license, and specs.
17 of the 25 Google vision models we track handle captioning: Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemma 4 12B, Gemini 3.5 Flash, Gemma 4 26B A4B, and Gemma 4 31B, and 11 more. Each model page lists its full task coverage, license, and specs.
Partly. 15 of the 25 Google vision models we track are open-weight (for example Gemma 4 12B, Gemma 4 26B A4B, and Gemma 4 31B), downloadable and self-hostable under their licenses (Apache 2.0, Custom, and MIT). The other 10 are proprietary and reached through Google's API.
We do not publish a Google-only ranking, so pick on constraints rather than a label. 18 Google models handle vision language; the most recent is Gemini 3.5 Flash-Lite (Jul 2026), and Gemma 4 12B is the newest open-weight option if you need to self-host. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.
We track 25 live Google vision models, 15 open-weight and 10 available through an API. The most recent addition is Gemini 3.5 Flash-Lite, released Jul 2026.
Yes. 15 of the 25 Google models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.
This page lists all 25 Google vision models in the Roboflow Playground catalog: 15 open-weight models you can self-host and 10 proprietary models accessed through an API. They cover vision language, captioning, and classification, among other tasks. 15 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.