Roboflow
Google

Google: Gemini 3.5 Flash

Gemini 3.5 Flash Overview

Gemini 3.5 Flash is a multimodal language model developed by Google DeepMind and released at Google I/O 2026. It is built on the Gemini 3 Flash reasoning foundation and introduces configurable thinking levels (minimal, low, medium, and high) that allow developers to tune the depth of internal reasoning before a response is generated. The model accepts text, image, video, audio, and PDF inputs and produces text output, with a 1 million token context window and up to 65,000 output tokens per request. It is natively multimodal, processing visual inputs alongside text to support tasks such as image captioning, classification, optical character recognition, object detection, and visual grounding, where the model references specific regions within an image or video frame.

Its vision capabilities extend to interpreting UI screenshots, diagrams, charts, and real-world scenes, as well as understanding video and live frame sequences for activity and scene recognition. The model supports combined tool use, including Google Search, URL context, code execution, and custom functions, within a single request, and it uses reasoning context from previous turns when thought signatures are present in the conversation history, enabling persistent multi-turn reasoning chains. Gemini 3.5 Flash carries a knowledge cutoff of January 2026 and is available via the Gemini API, Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform.

Gemini 3.5 Flash Interactive Demo

Gemini 3.5 Flash Details & Performance

Details

Resources

Vision Tasks

Chart Question AnsweringClassificationDocument Question AnsweringMulti-Label ClassificationOCRObject DetectionVisual Question Answering

Features

Multimodal VisionLLMs with Vision Capabilities

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Gemini 3.5 Flash Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated July 10, 2026Pricing updated July 21, 2026

Overall score#1 of 16
86.0%
Avg cost / sample#11 of 16
$0.0082
Avg speed / sample#6 of 16
4.8s
Avg tokens / sample
1.9K

Strengths and weaknesses

Gemini 3.5 Flash averages 86.0% across the six Vision Evals tasks, ranking #1 of 16 models overall.

It leads the field in Object Detection, Counting, and Identification.

It also places in the top three for Data Extraction and Reasoning.

Its weakest relative showing is OCR, ranking #6 of 16 at 91.1%.

At $0.0082 per sample it is the 11th cheapest of the 16 benchmarked models, and its average inference time of 4.8s per sample makes it the 6th fastest.

Performance profile

Field medianGemini 3.5 Flash

Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection
61.7%
#1 of 16$0.00965.9s
Counting
81.1%
#1 of 16$0.00753.9s
Identification
100.0%
#1 of 16$0.00422.7s
OCR
91.1%
#6 of 16$0.0167.4s
Data Extraction
94.8%
#2 of 16$0.00372.7s
Reasoning
87.0%
#2 of 16$0.00643.4s

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

16 models on the current benchmark · scores and efficiency pooled across all six tasks · Gemini 3.5 Flash highlighted

Gemini 3.5 Flash scores from a single evaluation run · Methodology

View all Vision Evals →

Gemini 3.5 Flash Pricing

Gemini 3.5 Flash costs $1.50 per 1M input tokens and $9.00 per 1M output tokens.

Input$1.50 / 1M tokens
Output$9.00 / 1M tokens
Cached input$0.150 / 1M tokens

Pricing updated Jul 21, 2026

Alternatives to Gemini 3.5 Flash

Other models worth comparing for similar use cases.

Google
Gemma 4 12B
Gemma 4 12B is an open-weight multimodal model from Google in the Gemma 4 family. It is intended for text and image understanding tasks such as visual question answering, OCR, captioning, and document understanding, with a smaller parameter footprint than the larger Gemma 4 variants.This entry is connected to Roboflow Playground vision evals for comparison. No runnable Playground workflow is configured yet, so the model page is used for discovery and benchmark context rather than direct hosted inference.
OpenAI
GPT-5 Mini
GPT-5 Mini, released by OpenAI on August 7, 2025, is a mid-tier variant of the GPT-5 family that balances cost, speed, and capability. It is multimodal, supporting both text and image inputs, and offers a substantial input context window of ~400,000 tokens with output lengths up to ~128,000 tokens. While less powerful than the full GPT-5, it inherits its safety tuning, instruction-following improvements, and multimodal reasoning, making it a practical choice for developers who need large context handling without the expense of premium models.GPT-5 Mini is optimized for affordability while retaining strong reasoning performance. Benchmarks show it outperforming earlier models such as GPT-4o on many multimodal and medical VQA tasks, though it lags behind GPT-5 on the most complex problems. Ideal use cases include prototyping, scalable content generation, document analysis, and mid-range reasoning tasks where efficiency and context capacity matter more than top-tier accuracy.
Anthropic
Claude Sonnet 5
Claude Sonnet 5 is a mid-tier large language model from Anthropic, released on June 30, 2026, as the latest model in the Sonnet series and a direct successor to Claude Sonnet 4.6. It is a hybrid reasoning model designed primarily for agentic workflows, software coding, and professional tasks. The model features a 1 million token context window, a 128k maximum output token limit, and runs adaptive thinking by default, giving API users fine-grained control over reasoning effort across five levels (low, medium, high, max, and extra-high). It uses an updated tokenizer shared with Opus 4.7 and later models, which produces approximately 30% more tokens for equivalent text compared to earlier Claude models. On benchmarks, Sonnet 5 scores 63.2% on agentic coding and 81.2% on OSWorld, narrowing the gap with Opus 4.8 while remaining at Sonnet-tier pricing.The model supports text and image input with text output, and accepts tools including browsers and terminals for autonomous multi-step task execution. Anthropic's safety evaluations report that Sonnet 5 shows a lower rate of undesirable behaviors than Sonnet 4.6 and is generally safer in agentic contexts, with improved resistance to prompt injection and reduced sycophancy. Cybersecurity safeguards equivalent to those on Opus 4.7 and 4.8 are active, though Anthropic notes the model was not deliberately trained on cybersecurity tasks. The model is proprietary and API-only, with no open weights.
Qwen
Qwen3.6 35B A3B
Qwen3.6-35B-A3B is a sparse Mixture-of-Experts (MoE) multimodal language model developed by the Qwen team at Alibaba Group. It carries 35 billion total parameters but activates only approximately 3 billion per forward pass via a learned routing mechanism, giving it the representational capacity of a large dense model at a fraction of the inference compute. The model is natively multimodal, processing images, documents, and video alongside text as a core architectural capability rather than an add-on. It supports a native context window of 262,144 tokens, extensible up to 1,010,000 tokens via YaRN. A key design feature is the unified thinking/non-thinking mode framework: users can switch between deliberate chain-of-thought reasoning and fast direct responses within a single model, and a "thinking preservation" option retains reasoning context across multi-turn agentic workflows to reduce redundant computation.The model is specifically optimized for agentic coding tasks, including repository-level reasoning, frontend workflow generation, multi-step tool use, and MCP (Model Context Protocol) integration. On SWE-bench Verified it scores 73.4%, on Terminal-Bench 2.0 it scores 51.5%, and on MCPMark it scores 37.0%. For vision-language tasks it achieves 92.0 on RefCOCO, 89.9 on OmniDocBench 1.5, and 83.7 on VideoMMMU. The model also supports Multi-Token Prediction (MTP) for speculative decoding. All Qwen3.6 open-weight models are released under the Apache 2.0 license.
OpenAI
GPT-5.6 Terra
GPT-5.6 Terra is the mid-tier reasoning model in OpenAI's GPT-5.6 family, which also includes the flagship Sol and the lightweight Luna. Introduced in a limited preview on June 26, 2026, and made broadly available on July 9, 2026, Terra accepts text and image input and produces text output, supporting vision, function calling, tool use, and agentic workflows. It is designed as a balanced option for everyday professional and production workloads — including coding assistance, document analysis, customer support, and multi-step agent tasks — where both output quality and cost efficiency matter. OpenAI positions Terra as delivering performance competitive with GPT-5.5 at approximately half the price, with a context window of around 1,050,000 tokens. On Terminal-Bench 2.1, Terra scores 84.3%, matching Claude Fable 5 on that benchmark. Under OpenAI's Preparedness Framework, Terra is rated High for cybersecurity and biological capabilities, meaning it demonstrates meaningful capability in those domains without reaching the Critical threshold.GPT-5.6 introduces a new naming convention in which the generation number (5.6) is paired with a durable capability tier name (Sol, Terra, or Luna), allowing each tier to advance on its own schedule. Terra carries the API identifier gpt-5.6-terra and supports the same reasoning effort controls available across the family, including adjustable reasoning depth. The model includes prompt caching with explicit cache breakpoints and a 30-minute minimum cache life, with cache writes billed at 1.25x the uncached input rate and cache reads receiving a 90% discount. GPT-5.6 Terra is a proprietary, closed-weights model served through the OpenAI API, Codex, and ChatGPT.

Other Google Gemini Flash models

Other versions in the same family as Gemini 3.5 Flash.

Gemini 3.5 Flash License

Proprietary

License terms and commercial-use guidance for Gemini 3.5 Flash.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Gemini 3.5 Flash Vision

Yes. Gemini 3.5 Flash accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Object Detection at 61.7% (#1 of 16). You can test it on your own image in the demo above.

Yes. its transcriptions match the ground truth 91.1% on average (#6 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 94.9%.

Yes. On Vision Evals, Gemini 3.5 Flash scores 61.7% mAP@50 on object detection (#1 of 16) and 81.1% exact-match accuracy on object counting.

On our benchmark's task mix, Gemini 3.5 Flash averages $0.0082 per sample at $1.50 per 1M input and $9.00 per 1M output tokens (#11 of 16 on cost), with an average speed of 4.8s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Gemini 3.5 Flash sits #1 of 16 at 86%, just ahead of Gemini 3.1 Pro (84.6%). See the full side-by-side: Gemini 3.5 Flash vs Gemini 3.1 Pro.