Gemma 4 12B is an open-weight multimodal model from Google in the Gemma 4 family. It is intended for text and image understanding tasks such as visual question answering, OCR, captioning, and document understanding, with a smaller parameter footprint than the larger Gemma 4 variants.
This entry is connected to Roboflow Playground vision evals for comparison. No runnable Playground workflow is configured yet, so the model page is used for discovery and benchmark context rather than direct hosted inference.
—
Usage
Past 30 DaysNot available
Not in Playground
Not yet ranked in arena
Gemma 4 12B has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Document Understanding | 8 / 9 | 88.9% |
| Object Understanding | 11 / 14 | 78.6% |
| Defect Detection | 11 / 15 | 73.3% |
| Spatial Understanding | 11 / 19 | 57.9% |
| Object Counting | 1 / 10 | 10% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
10 of 11 models plotted · 1 not yet evaluated
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| Claude Opus 4.8 | 67.2% | 2.2K | $0.012 | Compare |
| Claude Opus 4.7 | 67.2% | 2.6K | $0.015 | Compare |
| Gemma 4 31B | 67.2% | 467 | $0.0001 | Compare |
| Claude Opus 4.6 | 64.2% | 2.3K | $0.014 | Compare |
| GPT-5.4 Nano | 62.7% | 1.8K | $0.0004 | Compare |
| Gemma 4 12B(this model) | 62.7% | — | — | — |
| Llama 4 Maverick | 59.7% | 2.4K | $0.0005 | Compare |
| Claude Sonnet 4.5 | 59.7% | 2.3K | $0.0092 | Compare |
| Claude Opus 4.1 | 59.7% | 2.1K | $0.040 | Compare |
| Claude Haiku 4.5 | 58.2% | 2.3K | $0.0030 | Compare |
| GPT-5 Nano | 58.2% | 2.7K | $0.0003 | Compare |
Other models worth comparing for similar use cases.
License terms and commercial-use guidance for Gemma 4 12B.
This model is released under the Apache License 2.0, a permissive open-source license that allows commercial use, modification, distribution, and patent use.
Yes. Under the terms of the Apache 2.0 license, you can freely use this model for commercial purposes, including in proprietary products. You must retain the copyright notice and disclaimers when redistributing.
License information is provided as a guide and is not legal advice.
Yes. Gemma 4 12B accepts image input, and on Roboflow's previous vision benchmark it passed 62.7% of visual understanding tasks (#40 of 77).
Gemma 4 12B has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.