Roboflow
Z.ai

GLM-4.6V Flash

Benchmark-only page (local)

GLM-4.6V Flash has no model page of its own yet — this page exists so its self-hosted GPU benchmark results have somewhere to live. See how it scores against every other model on Vision Evals.

Family
glm-4.6v
Architecture
dense
Quantizations
BF16, AWQ INT4
Weights on disk
8.9–20.6 GB

Self-hosted benchmarks

2 quantizations of GLM-4.6V Flash, served with vLLM from the published weights and scored on the same tasks as the hosted models. Expand a row for the GPUs it was measured on.

Self-hosted rows run with thinking on, the same setting as the hosted frontier models. Rows marked with a run count are the mean of three runs per task under the benchmark protocol; the rest are single runs awaiting their re-run. Some hosted open-weight rows still run without thinking, so a self-hosted quant can score above its own hosted API. Quantization still costs a little precision, and scores vary between runs.

QuantizationWeightsOverall scoreDetection mAP50Single-stream tok/s
BF163 runsRanked
20.6 GB
59.7%
35.0%120.3
8.9 GB
56.4%
42.5%209.1