Roboflow

Google Vision Models

Compare all 25 Google vision models we track, 15 of them open-weight. Try 15 of them on your own images, free in the Roboflow Playground.

25 models · 15 open-weight · 15 free to try · Updated Jul 2026

About Google Vision Models

Google ships vision models on two tiers, and our catalog tracks 25 of them. 15 are open-weight: the weights are downloadable, so you can self-host under their own licenses. The other 10 are proprietary, reached through an API and billed by the provider. The open-weight side runs on Apache 2.0, Custom, and MIT licenses. The newest additions are Gemini 3.5 Flash-Lite (Jul 2026) and Gemini 3.6 Flash (Jul 2026). The lineup reaches back to Google Vision OCR, released Feb 2016.

Between them, the Google models we list cover 17 distinct vision tasks. The widest coverage is vision language (18 models), captioning (17), and classification (17). For vision language, start with Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, or Gemma 4 12B.

On the open-weight tier, Gemma 4 31B is the largest at 31B parameters and MobileNetV2 the smallest at ~3.4M, small enough to run on a single GPU or on-device. 10 of the 15 open Google models carry a permissive license, so commercial use is straightforward; the other 5 ship under terms worth reading before you deploy.

On the API tier, Google's most recent entry is Gemini 3.5 Flash-Lite (Jul 2026), billed per token by the provider rather than run on your own hardware. Which tier to start on is a constraints question, not a quality one: reach for the open-weight side when you need offline inference, predictable per-image cost, or a checkpoint you can fine-tune on your own data, and for the API side when you want broad general reasoning without managing GPUs. 15 of the 25 Google vision models run live in the Roboflow Playground, so you can run the same image through several of them and compare the answers before committing to one.

Which Google Model Should You Use?

What each of the 25 Google vision models in our catalog is built for, and how you run it.

Gemini 3.5 Flash-Lite
Best for: Object Detection and Video Classification. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 3.6 Flash
Best for: Object Detection and Video Classification. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemma 4 12B
Best for: OCR and Captioning. Open weights: Apache 2.0 license, 12B parameters. Self-host it or run it here.
Gemini 3.5 Flash
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemma 4 26B A4B
Best for: Object Detection and OCR. Open weights: Apache 2.0 license, 25.2B parameters. Self-host it or run it here. Runnable in the Playground.
Gemma 4 31B
Best for: Object Detection and OCR. Open weights: Apache 2.0 license, 31B parameters. Self-host it or run it here. Runnable in the Playground.
Gemini 3.1 Flash-Lite
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 3.1 Pro
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 3 Flash
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 2.5 Flash-Lite
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 2.5 Flash
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemini 2.5 Pro
Best for: Object Detection and OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.
Gemma 3 12B
Best for: OCR and Document Question Answering. Open weights: Custom license, 12B parameters. Self-host it or run it here. Runnable in the Playground.
Gemma 3 27B
Best for: OCR and Document Question Answering. Open weights: Custom license. Self-host it or run it here. Runnable in the Playground.
Gemma 3 4B
Best for: OCR and Document Question Answering. Open weights: Custom license, 4B parameters. Self-host it or run it here. Small enough for on-device and edge deployment. Runnable in the Playground.
PaliGemma 2
Best for: OCR and Captioning. Open weights: Custom license, 3B, 10B, 28B parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
PaliGemma
Best for: Captioning and Visual Question Answering. Open weights: Custom license, 3B parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
SigLIP
Best for: Image Embedding and Image Similarity. Open weights: Apache 2.0 license, 200M-900M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
OWL-ViT
Best for: Open Vocabulary Object Detection and Object Detection. Open weights: Apache 2.0 license. Self-host it or run it here.
Vision Transformer (ViT)
Best for: Classification. Open weights: Apache 2.0 license, 86M-632M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
EfficientDet
Best for: Object Detection. Open weights: Apache 2.0 license, 3.9M-51.9M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
MediaPipe
Best for: Semantic Segmentation and Object Detection. Open weights: Apache 2.0 license. Self-host it or run it here.
MobileNet SSD v2
Best for: Object Detection. Open weights: MIT license, 15.3M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
MobileNetV2
Best for: Classification. Open weights: Apache 2.0 license, ~3.4M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.
Google Vision OCR
Best for: OCR. Proprietary: no downloadable weights, billed per token through Google's API. Runnable in the Playground.

Open-Source Google Models

15 models with downloadable weights you can self-host under their licenses (Apache 2.0, Custom, and MIT). 5 run live in the Playground through hosted APIs, so self-hosting is optional.

Google
12BApache 2.0Jun 2026
Google
25.2BApache 2.0Apr 2026
Try
Google
31BApache 2.0Apr 2026
Try
Google
12BCustomMar 2025
Google
CustomMar 2025
Google
4BCustomMar 2025
Google
3B, 10B, 28BCustomDec 2024
Google
PaliGemmaGoogle
3BCustomMay 2024

Google Models via API

10 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.

$0.30$2.501.0MJul 2026
Try
Google
$1.50$7.501MJul 2026
Try
Google
$1.50$9.001.0MMay 2026
Try
$0.25$1.501MMar 2026
Try
Google
$2.00$12.001MFeb 2026
Try
Google
$0.50$3.001MDec 2025
Try
$0.10$0.401MJul 2025
Try
Google
$0.30$2.501MJul 2025
Try

Frequently Asked Questions About Google Vision Models

Which Google models can do vision language?

18 of the 25 Google vision models we track handle vision language: Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemma 4 12B, Gemini 3.5 Flash, Gemma 4 26B A4B, and Gemma 4 31B, and 12 more. Each model page lists its full task coverage, license, and specs.

Which Google models can do captioning?

17 of the 25 Google vision models we track handle captioning: Gemini 3.5 Flash-Lite, Gemini 3.6 Flash, Gemma 4 12B, Gemini 3.5 Flash, Gemma 4 26B A4B, and Gemma 4 31B, and 11 more. Each model page lists its full task coverage, license, and specs.

Are Google vision models open source?

Partly. 15 of the 25 Google vision models we track are open-weight (for example Gemma 4 12B, Gemma 4 26B A4B, and Gemma 4 31B), downloadable and self-hostable under their licenses (Apache 2.0, Custom, and MIT). The other 10 are proprietary and reached through Google's API.

What is the best Google model for vision language?

We do not publish a Google-only ranking, so pick on constraints rather than a label. 18 Google models handle vision language; the most recent is Gemini 3.5 Flash-Lite (Jul 2026), and Gemma 4 12B is the newest open-weight option if you need to self-host. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.

How many Google vision models are on Roboflow Playground?

We track 25 live Google vision models, 15 open-weight and 10 available through an API. The most recent addition is Gemini 3.5 Flash-Lite, released Jul 2026.

Can I try Google vision models for free?

Yes. 15 of the 25 Google models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.

This page lists all 25 Google vision models in the Roboflow Playground catalog: 15 open-weight models you can self-host and 10 proprietary models accessed through an API. They cover vision language, captioning, and classification, among other tasks. 15 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.