Gemma 4 31B is the largest dense model in Google's Gemma 4 family, built from the same research as Gemini 3 and released as open weights under the Apache 2.0 license. It supports a 256K token context window with text and image input, configurable thinking mode for step-by-step reasoning, and multilingual support across 140+ languages. The unquantized model fits on a single 80GB GPU.
For vision tasks, Gemma 4 31B supports image understanding with variable aspect ratios and resolutions, and can output structured bounding boxes for UI element detection, making it useful for document parsing and UI understanding. Compared to Gemma 3, it delivers stronger reasoning and multimodal performance. It is part of a four-size family alongside the 26B A4B MoE variant and two on-device models (E2B, E4B), with the 31B dense variant optimized for output quality and fine-tuning over inference speed.
Drag and drop an image here, or click to browse
Captioning will run automatically
—
Usage
Past 30 DaysGemma 4 31B has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Document Understanding | 8 / 9 | 88.9% |
| Defect Detection | 12 / 15 | 80% |
| Spatial Understanding | 14 / 19 | 73.7% |
| Object Understanding | 10 / 14 | 71.4% |
| Object Counting | 1 / 10 | 10% |
| Category | Passed | Score |
|---|---|---|
| License Plate Recognition | 28 / 30 | 93.3% |
| Focused Scene OCR | 86 / 99 | 86.9% |
| VQA & Extraction | 51 / 60 | 85% |
| Text Recognition | 24 / 30 | 80% |
| Handwritten Math | 5 / 10 | 50% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →Gemma 4 31B costs $0.120 per 1M input tokens and $0.370 per 1M output tokens.
Pricing updated Jul 21, 2026
Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
11 of 11 models plotted
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| Gemini 3.1 Flash-Lite | 68.7% | 1.1K | $0.0003 | Compare |
| Gemma 4 26B A4B | 68.7% | 531 | $0.0001 | Compare |
| Qwen3.6 Plus | 68.7% | 1.6K | $0.0005 | Compare |
| Claude Opus 4.8 | 67.2% | 2.2K | $0.012 | Compare |
| Claude Opus 4.7 | 67.2% | 2.6K | $0.015 | Compare |
| Gemma 4 31B(this model) | 67.2% | 467 | $0.0001 | — |
| Claude Opus 4.6 | 64.2% | 2.3K | $0.014 | Compare |
| GPT-5.4 Nano | 62.7% | 1.8K | $0.0004 | Compare |
| Llama 4 Maverick | 59.7% | 2.4K | $0.0005 | Compare |
| Claude Sonnet 4.5 | 59.7% | 2.3K | $0.0092 | Compare |
| Claude Opus 4.1 | 59.7% | 2.1K | $0.040 | Compare |
Other models worth comparing for similar use cases.
License terms and commercial-use guidance for Gemma 4 31B.
This model is released under the Apache License 2.0, a permissive open-source license that allows commercial use, modification, distribution, and patent use.
Yes. Under the terms of the Apache 2.0 license, you can freely use this model for commercial purposes, including in proprietary products. You must retain the copyright notice and disclaimers when redistributing.
License information is provided as a guide and is not legal advice.
Yes. Gemma 4 31B accepts image input, and on Roboflow's previous vision benchmark it passed 67.2% of visual understanding tasks (#30 of 77) and scored 84.7% on OCR. You can test it on your own image in the demo above.
Gemma 4 31B has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.
Yes. The demo on this page runs Gemma 4 31B in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.