Roboflow

Vision Evals

See which AI vision models are best at object detection, counting, identification, OCR, data extraction, and reasoning. Every model runs the same real-world tasks and is scored against ground truth by Roboflow.

Looking for rankings based on real user votes? See Arena Rankings

16 models evaluated|6 tasks

Evals updated July 10, 2026Pricing updated July 21, 2026

Overall ranking

Score key:≥75%40–74%<40%
1
86.0%
1.9K$0.00824.8s
2
84.6%
1.5K$0.00685.9s
3
79.6%
2.0K$0.0298.4s
4
77.2%
1.5K$0.00173.9s
5
74.6%
2.4K$0.0239.9s
6
74.1%
2.3K$0.00455.2s
7
73.5%
2.3K$0.0106.0s
8
OpenAIGPT-5.5
71.8%
2.2K$0.02110.4s
9
67.9%
654$0.00364.7s
10
66.4%
1.6K$0.00078.2s
11
QwenQwen 3.7 Plus
66.2%
1.4K$0.00065.8s
12
Z.aiGLM 5V Turbo
65.9%
1.9K$0.00277.1s
13
65.6%
1.9K$0.00534.0s
14
65.3%
1.8K$0.00264.5s
15
64.8%
1.9K$0.0134.2s
16
MoonshotAIKimi K2.6
57.0%
1.9K$0.002110.1s

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

16 models on the current benchmark · scores and efficiency pooled across all six tasks

Results by task

Score key:≥75%40–74%<40%
ModelObject DetectionCountingIdentificationOCRData ExtractionReasoning
61.781.1100.091.194.887.0
56.971.6100.092.694.891.3
40.663.5100.094.092.887.0
39.367.693.887.696.978.3
46.273.081.390.782.573.9
43.366.278.188.481.487.0
44.767.678.188.879.482.6
OpenAIGPT-5.5
13.864.990.691.287.682.6
26.452.793.888.884.560.9
42.347.390.688.186.643.5
QwenQwen 3.7 Plus
49.250.084.486.583.543.5
Z.aiGLM 5V Turbo
48.248.684.489.381.443.5
18.056.881.391.789.756.5
3.960.878.188.182.578.3
18.652.775.093.887.660.9
MoonshotAIKimi K2.6
32.250.062.588.578.330.4

Cell color follows the absolute score, on the same scale as the score key above. The column leader is bold. Hover a cell for the full breakdown.

What are Vision Evals?

Vision Evals measures what vision language models can actually do with real images. Every model receives the same samples across six tasks: Object Detection, Counting, Identification, OCR, Data Extraction, and Reasoning. Answers are scored against ground truth, so the scores are directly comparable. No human votes, no subjective judgment.

Methodology

Each model runs the full sample set for every task in a single evaluation pass at low reasoning effort. Object Detection is scored by mean Average Precision (mAP@50 as the headline, with mAP@75 and mAP@50:95 reported on the task page), OCR by mean similarity to the ground-truth transcription, and the other four tasks by exact-match accuracy. The overall Average column is the unweighted mean of the six task scores.

Token usage is measured directly from each provider’s API response, and output counts include reasoning tokens. Estimated cost per sample multiplies that measured input and output usage by the model’s published per-1M pricing at the time of our last price sync, so prices stay current without re-running the benchmark; it is an estimate on this benchmark, not a universal model cost. Speed is wall-clock inference time per sample.

Last evaluated: July 10, 2026

Looking for the previous benchmark?

The previous version of Vision Evals is preserved with its full leaderboards. It covers many more models, including deprecated ones that can no longer be re-run. See Vision Evals (legacy)

Frequently Asked Questions

Arena Rankings show which model people prefer when comparing two outputs side by side. Vision Evals shows which model actually gets the right answer on specific tasks, scored against ground truth. A model can rank highly in the Arena but still struggle with counting or detection. Use both together to make a better decision.

Each of the six tasks produces one headline score: mAP@50 for Object Detection, mean similarity for OCR, and exact-match accuracy for the rest. The Average is the unweighted mean of those six scores, so every task counts equally.

The previous version of Vision Evals is preserved at its own page and still covers many models that predate this benchmark, including deprecated models that can no longer be re-run. You can browse it at /evals/legacy.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.