Gemini 3 Flash is a proprietary multimodal large language model developed by Google through Google DeepMind, designed to deliver fast, cost-efficient reasoning across real-time products and developer workflows. Released in December 2025, it is the Flash-tier variant of the Gemini 3 family, balancing low latency with reasoning quality approaching Pro models.
The model supports text, images, audio, and video, with an exceptionally large context window of roughly one million input tokens and outputs up to ~65k tokens. It emphasizes rapid responses for coding, summarization, analysis, and agentic tasks, and exposes configurable “thinking levels” via API to trade speed for deeper reasoning. Today, Gemini 3 Flash positions itself as a high-throughput, production-ready model, serving as the default in the Gemini app and Google Search’s AI Mode, optimized for scalable, interactive AI applications.
Drag and drop an image here, or click to browse
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated July 10, 2026Pricing updated July 21, 2026
Gemini 3 Flash averages 77.2% across the six Vision Evals tasks, ranking #4 of 16 models overall.
It leads the field in Data Extraction.
Its weakest relative showing is OCR, ranking #15 of 16 at 87.6%.
At $0.0017 per sample it is the 3rd cheapest of the 16 benchmarked models, and its average inference time of 3.9s per sample makes it the fastest.
Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 39.3% | #10 of 16 | $0.0022 | 5.1s | |
| Counting | 67.6% | #4 of 16 | $0.0012 | 2.9s | |
| Identification | 93.8% | #4 of 16 | $0.0009 | 2.2s | |
| OCR | 87.6% | #15 of 16 | $0.0024 | 4.2s | |
| Data Extraction | 96.9% | #1 of 16 | $0.0008 | 2.0s | |
| Reasoning | 78.3% | #7 of 16 | $0.0017 | 3.6s |
Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
16 models on the current benchmark · scores and efficiency pooled across all six tasks · Gemini 3 Flash highlighted
Gemini 3 Flash scores from a single evaluation run · Methodology
View all Vision Evals →Gemini 3 Flash costs $0.500 per 1M input tokens and $3.00 per 1M output tokens.
Pricing updated Jul 21, 2026
Other models worth comparing for similar use cases.
Other versions in the same family as Gemini 3 Flash.
License terms and commercial-use guidance for Gemini 3 Flash.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. Gemini 3 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Data Extraction at 96.9% (#1 of 16). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 87.6% on average (#15 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 96.9%.
Not its strength. On Vision Evals, Gemini 3 Flash scores 39.3% mAP@50 on object detection (#10 of 16) and 67.6% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, Gemini 3 Flash averages $0.0017 per sample at $0.50 per 1M input and $3.00 per 1M output tokens (#3 of 16 on cost), with an average speed of 3.9s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Gemini 3 Flash sits #4 of 16 at 77.2%, just behind Claude Fable 5 (79.6%) and just ahead of GPT-5.6 Sol (74.6%). See the full side-by-side: Gemini 3 Flash vs Claude Fable 5.