Qwen3.5 4B has no model page of its own yet — this page exists so its self-hosted GPU benchmark results have somewhere to live. See how it scores against every other model on Vision Evals.
One quantization of Qwen3.5 4B, served with vLLM from the published weights and scored on the same tasks as the hosted models. Expand a row for the GPUs it was measured on.
Self-hosted rows run with thinking on, the same setting as the hosted frontier models. Rows marked with a run count are the mean of three runs per task under the benchmark protocol; the rest are single runs awaiting their re-run. Some hosted open-weight rows still run without thinking, so a self-hosted quant can score above its own hosted API. Quantization still costs a little precision, and scores vary between runs.
| Quantization | Weights | Overall score | Detection mAP50 | Single-stream tok/s |
|---|---|---|---|---|
| 9.3 GB | 58.2% | 38.1% | 200.6 |