Gemma 3 4B, released on March 12, 2025, is the mid-sized member of Google DeepMind’s open-weight Gemma 3 family. With about 4 billion parameters, it is multimodal—supporting text and image inputs and generating text outputs. Like the larger Gemma 3 models, it features a 128,000-token input context window with an output capacity of ~8,192 tokens, enabling it to handle long documents and mixed text–image reasoning tasks.
The 4B variant is designed as a balance between efficiency and capability: it offers multilingual support across 140+ languages, strong summarization and reasoning performance, and compatibility with moderate hardware. Inference can run with ~6.4 GB VRAM in BF16, or significantly less in quantized 8-bit (~4.4 GB) or 4-bit (~3.4 GB) modes, making it accessible to developers outside large-scale infrastructure. While it lags behind the 12B and 27B versions on the most complex reasoning and multimodal benchmarks, its lower compute footprint makes it ideal for research, prototyping, and practical deployment where efficiency matters.
Drag and drop an image here, or click to browse
Captioning will run automatically
—
Usage
Past 30 DaysGemma 3 4B has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Defect Detection | 9 / 15 | 60% |
| Document Understanding | 5 / 9 | 55.6% |
| Object Understanding | 6 / 14 | 42.9% |
| Spatial Understanding | 5 / 19 | 26.3% |
| Object Counting | 0 / 10 | 0% |
| Category | Passed | Score |
|---|---|---|
| License Plate Recognition | 26 / 30 | 86.7% |
| Text Recognition | 22 / 30 | 73.3% |
| Focused Scene OCR | 63 / 99 | 63.6% |
| VQA & Extraction | 35 / 60 | 58.3% |
| Handwritten Math | 1 / 10 | 10% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →Gemma 3 4B costs $0.050 per 1M input tokens and $0.100 per 1M output tokens.
Pricing updated Aug 12, 2026
Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
6 of 7 models plotted · 1 not yet evaluated
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| Claude Haiku 4.5 | 58.2% | 2.3K | $0.0030 | Compare |
| GPT-5 Nano | 58.2% | 2.7K | $0.0003 | Compare |
| Qwen3.5 397B A17B | 58.2% | 1.5K | $0.0008 | Compare |
| Gemini 2.5 Flash | 55.2% | 476 | $0.0005 | Compare |
| Gemini 2.5 Flash-Lite | 53.7% | 301 | <$0.0001 | Compare |
| Gemma 3 4B(this model) | 37.3% | — | — | — |
| Kimi K2.5 | 35.8% | 2.7K | $0.0031 | Compare |
Other models worth comparing for similar use cases.
Gemma 3 4B ships under a custom, model-specific license rather than a standard permissive or restrictive one, so the Gemma 3 4B license has to be read directly. Custom model licenses range from effectively permissive to research-only.
Uncertainty around licensing can delay or stop a project, and acceptable-use policies attached to custom licenses are binding terms rather than guidance. Review them alongside the Gemma 3 4B license before production deployment.
If the custom terms rule out your use case, a commercial license from the rights holder is the way through. Roboflow's licensing page lists the supported models whose commercial license is included in a Roboflow plan, so it is worth checking whether Gemma 3 4B — or a permissively licensed alternative — fits your deployment.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is released under a custom license that does not match a standard open-source identifier. Read the full license text linked from the model documentation.
Custom licenses vary widely in what they permit. Many model-specific custom licenses include commercial-use restrictions (e.g., non-commercial weights, named-user limits, or jurisdiction restrictions). Read the full license before deploying commercially.
Custom licenses are model-specific. Always check the per-model License Notes section above and the linked official license text.
License information is provided as a guide and is not legal advice.
Yes. Gemma 3 4B accepts image input, and on Roboflow's previous vision benchmark it passed 37.3% of visual understanding tasks (#73 of 77) and scored 64.2% on OCR. You can test it on your own image in the demo above.
Gemma 3 4B has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.
Yes. The demo on this page runs Gemma 3 4B in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.