Roboflow

Vision Evals

See which AI vision models are best at object detection, counting, identification, OCR, data extraction, and reasoning. Every model runs the same real-world tasks and is scored against ground truth by Roboflow.

Looking for rankings based on real user votes? See Arena Rankings

28 models evaluated|6 tasks

Evals updated August 12, 2026Pricing updated August 13, 2026

Overall ranking

Score key:≥75%40–74%<40%
1
86.6%
2.1K$0.0115.82s
2
84.0%
2.1K$0.007418.02s
3
83.1%
1.7K$0.00937.81s
4
83.1%
1.8K$0.00634.73s
5
80.4%
2.9K$0.00717.78s
6
79.3%
3.0K$0.006911.40s
7
78.8%
2.2K$0.0348.71s
8
AnthropicClaude Opus 5
77.1%
2.1K$0.0177.38s
9
76.9%
2.2K$0.02511.72s
10
74.9%
1.7K$0.00214.10s
11
OpenAIGPT-5.5
73.9%
2.0K$0.0229.35s
12
72.4%
2.1K$0.00447.15s
13
71.5%
2.1K$0.00056.55s
14
70.8%
2.3K$0.00138.70s
15
69.6%
1.6K$0.00142.70s
16
GrokGrok 4.6NEW
67.8%
2.1K$0.00697.39s
17
QwenQwen 3.7 Plus
67.4%
1.5K$0.00087.01s
18
66.8%
2.1K$0.0165.20s
19
MoonshotAIKimi K3
66.5%
2.6K$0.01112.71s
20
66.4%
2.1K$0.00644.84s
21
66.0%
796$0.00506.11s
22
65.8%
1.9K$0.00069.17s
23
Z.aiGLM 5V Turbo
65.3%
2.2K$0.00316.35s
24
64.3%
2.5K$0.007714.33s
25
64.3%
1.7K$0.00077.38s
26
63.5%
1.9K$0.00305.35s
27
61.7%
1.5K$0.00016.32s
28
MoonshotAIKimi K2.6
59.0%
2.3K$0.003229.57s

Score vs. cost

Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

28 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort

Results by task

Score key:≥75%40–74%<40%
ModelObject DetectionCountingIdentificationOCRData ExtractionReasoning
68.781.1100.091.194.884.1
77.182.490.692.887.673.5
67.471.6100.092.694.872.2
56.082.496.988.494.880.1
60.274.390.693.888.774.8
58.475.787.592.586.674.8
56.463.5100.094.092.866.2
AnthropicClaude Opus 5
54.470.384.493.288.771.5
68.273.081.390.782.565.6
38.667.693.887.696.964.9
OpenAIGPT-5.5
41.764.990.691.287.667.5
60.767.678.188.879.459.6
59.966.278.188.481.455.0
41.066.281.392.186.657.6
57.552.781.387.490.748.3
GrokGrok 4.6NEW
20.270.378.192.084.561.6
QwenQwen 3.7 Plus
60.150.084.486.583.539.7
38.652.775.093.887.653.0
MoonshotAIKimi K3
51.946.081.393.084.542.4
36.156.881.391.789.743.0
33.752.793.888.884.542.4
52.147.390.688.186.629.8
Z.aiGLM 5V Turbo
56.548.684.489.381.431.8
18.055.478.192.583.558.3
58.854.078.184.578.331.8
16.160.878.188.182.555.6
42.846.084.484.178.334.4
MoonshotAIKimi K2.6
35.750.062.588.578.339.1

Cell color follows the absolute score, on the same scale as the score key above. The column leader is bold. Hover a cell for the full breakdown.

What are Vision Evals?

Vision Evals measures what vision language models can actually do with real images. Every model receives the same samples across six tasks: Object Detection, Counting, Identification, OCR, Data Extraction, and Reasoning. Answers are scored against ground truth, so the scores are directly comparable. No human votes, no subjective judgment.

Methodology

Each model runs the full sample set for every task in a single evaluation pass at low reasoning effort. Reasoning is additionally run at high effort, and the Reasoning leaderboard ranks both passes side by side; every other table on this site, including the Average column, uses the low-effort pass so all six tasks stay comparable. Object Detection is scored by mean Average Precision (mAP@50 as the headline, with mAP@75 and mAP@50:95 reported on the task page), OCR by mean similarity to the ground-truth transcription, and the other four tasks by exact-match accuracy. The overall Average column is the unweighted mean of the six task scores.

Token usage is measured directly from each provider’s API response, and output counts include reasoning tokens. Estimated cost per sample multiplies that measured input and output usage by the model’s published per-1M pricing at the time of our last price sync, so prices stay current without re-running the benchmark; it is an estimate on this benchmark, not a universal model cost. Speed is wall-clock inference time per sample.

Last evaluated: August 12, 2026

Looking for the previous benchmark?

The previous version of Vision Evals is preserved with its full leaderboards. It covers many more models, including deprecated ones that can no longer be re-run. See Vision Evals (legacy)

Frequently Asked Questions

Arena Rankings show which model people prefer when comparing two outputs side by side. Vision Evals shows which model actually gets the right answer on specific tasks, scored against ground truth. A model can rank highly in the Arena but still struggle with counting or detection. Use both together to make a better decision.

Each of the six tasks produces one headline score: mAP@50 for Object Detection, mean similarity for OCR, and exact-match accuracy for the rest. The Average is the unweighted mean of those six scores, so every task counts equally.

The previous version of Vision Evals is preserved at its own page and still covers many models that predate this benchmark, including deprecated models that can no longer be re-run. You can browse it at /evals/legacy.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.