Gemini 3.5 Flash is a multimodal language model developed by Google DeepMind and released at Google I/O 2026. It is built on the Gemini 3 Flash reasoning foundation and introduces configurable thinking levels (minimal, low, medium, and high) that allow developers to tune the depth of internal reasoning before a response is generated. The model accepts text, image, video, audio, and PDF inputs and produces text output, with a 1 million token context window and up to 65,000 output tokens per request. It is natively multimodal, processing visual inputs alongside text to support tasks such as image captioning, classification, optical character recognition, object detection, and visual grounding, where the model references specific regions within an image or video frame.
Its vision capabilities extend to interpreting UI screenshots, diagrams, charts, and real-world scenes, as well as understanding video and live frame sequences for activity and scene recognition. The model supports combined tool use, including Google Search, URL context, code execution, and custom functions, within a single request, and it uses reasoning context from previous turns when thought signatures are present in the conversation history, enabling persistent multi-turn reasoning chains. Gemini 3.5 Flash carries a knowledge cutoff of January 2026 and is available via the Gemini API, Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform.
Drag and drop an image here, or click to browse
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated July 10, 2026Pricing updated July 21, 2026
Gemini 3.5 Flash averages 86.0% across the six Vision Evals tasks, ranking #1 of 16 models overall.
It leads the field in Object Detection, Counting, and Identification.
It also places in the top three for Data Extraction and Reasoning.
Its weakest relative showing is OCR, ranking #6 of 16 at 91.1%.
At $0.0082 per sample it is the 11th cheapest of the 16 benchmarked models, and its average inference time of 4.8s per sample makes it the 6th fastest.
Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 61.7% | #1 of 16 | $0.0096 | 5.9s | |
| Counting | 81.1% | #1 of 16 | $0.0075 | 3.9s | |
| Identification | 100.0% | #1 of 16 | $0.0042 | 2.7s | |
| OCR | 91.1% | #6 of 16 | $0.016 | 7.4s | |
| Data Extraction | 94.8% | #2 of 16 | $0.0037 | 2.7s | |
| Reasoning | 87.0% | #2 of 16 | $0.0064 | 3.4s |
Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
16 models on the current benchmark · scores and efficiency pooled across all six tasks · Gemini 3.5 Flash highlighted
Gemini 3.5 Flash scores from a single evaluation run · Methodology
View all Vision Evals →Gemini 3.5 Flash costs $1.50 per 1M input tokens and $9.00 per 1M output tokens.
Pricing updated Jul 21, 2026
Other models worth comparing for similar use cases.
Other versions in the same family as Gemini 3.5 Flash.
License terms and commercial-use guidance for Gemini 3.5 Flash.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. Gemini 3.5 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Object Detection at 61.7% (#1 of 16). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 91.1% on average (#6 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 94.9%.
Yes. On Vision Evals, Gemini 3.5 Flash scores 61.7% mAP@50 on object detection (#1 of 16) and 81.1% exact-match accuracy on object counting.
On our benchmark's task mix, Gemini 3.5 Flash averages $0.0082 per sample at $1.50 per 1M input and $9.00 per 1M output tokens (#11 of 16 on cost), with an average speed of 4.8s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Gemini 3.5 Flash sits #1 of 16 at 86%, just ahead of Gemini 3.1 Pro (84.6%). See the full side-by-side: Gemini 3.5 Flash vs Gemini 3.1 Pro.