Roboflow
Google

Gemma 4 E4B

Benchmark-only page (local)

Gemma 4 E4B has no model page of its own yet — this page exists so its self-hosted GPU benchmark results have somewhere to live. See how it scores against every other model on Vision Evals.

Family
gemma-4
Architecture
dense
Quantizations
BF16, QAT W4A16
Weights on disk
11.5–16.0 GB

Self-hosted benchmarks

2 quantizations of Gemma 4 E4B, served with vLLM from the published weights and scored on the same tasks as the hosted models. Expand a row for the GPUs it was measured on.

Self-hosted rows run with thinking on, the same setting as the hosted frontier models. Rows marked with a run count are the mean of three runs per task under the benchmark protocol; the rest are single runs awaiting their re-run. Some hosted open-weight rows still run without thinking, so a self-hosted quant can score above its own hosted API. Quantization still costs a little precision, and scores vary between runs.

QuantizationWeightsOverall scoreDetection mAP50Single-stream tok/s
BF163 runsRanked
16.0 GB
47.6%
23.8%163.7
11.5 GB
38.7%
13.7%223.6