Gemma 4 E2B has no model page of its own yet — this page exists so its self-hosted GPU benchmark results have somewhere to live. See how it scores against every other model on Vision Evals.
One quantization of Gemma 4 E2B, served with vLLM from the published weights and scored on the same tasks as the hosted models. Expand a row for the GPUs it was measured on.
Self-hosted rows run with thinking on, the same setting as the hosted frontier models. Rows marked with a run count are the mean of three runs per task under the benchmark protocol; the rest are single runs awaiting their re-run. Some hosted open-weight rows still run without thinking, so a self-hosted quant can score above its own hosted API. Quantization still costs a little precision, and scores vary between runs.
| Quantization | Weights | Overall score | Detection mAP50 | Single-stream tok/s |
|---|---|---|---|---|
| 10.2 GB | 43.9% | 19.0% | 258.1 |