Roboflow

Molmo2 8B

Benchmark-only page (local)

Molmo2 8B has no model page of its own yet — this page exists so its self-hosted GPU benchmark results have somewhere to live. See how it scores against every other model on Vision Evals.

Family
molmo2
Architecture
dense
Quantizations
BF16, FP8
Weights on disk
13.8–34.6 GB
Hugging Face
allenai/Molmo2-8B

Self-hosted benchmarks

2 quantizations of Molmo2 8B, served with vLLM from the published weights and scored on the same tasks as the hosted models. Expand a row for the GPUs it was measured on.

Self-hosted rows run with thinking on, the same setting as the hosted frontier models. Rows marked with a run count are the mean of three runs per task under the benchmark protocol; the rest are single runs awaiting their re-run. Some hosted open-weight rows still run without thinking, so a self-hosted quant can score above its own hosted API. Quantization still costs a little precision, and scores vary between runs.

QuantizationWeightsOverall scoreDetection mAP50Single-stream tok/s
BF16no thinking modeRanked
34.6 GB
40.9%
0.0%139.2
FP8no thinking mode
13.8 GB
37.2%
0.0%227.5