Roboflow

Vision Evals

See which AI vision models are best at object detection, counting, identification, OCR, data extraction, and reasoning. Every model runs the same real-world tasks and is scored against ground truth by Roboflow.

57 models evaluated|6 tasks

Evals updated September 22, 2026Pricing updated September 22, 2026

Overall ranking

Score key:≥75%40–74%<40%
1
86.6%
1.9K$0.0306.67s
2
86.0%
2.1K$0.01114.77s
3
AnthropicClaude Opus 5.5NEW
85.5%
2.2K$0.01412.76s
4
85.2%
1.8K$0.003116.50s
5
85.1%
1.8K$0.003311.65s
6
83.9%
2.1K$0.007417.25s
7
83.3%
1.7K$0.00937.81s
8
83.0%
1.8K$0.003214.66s
9
81.3%
2.2K$0.0358.28s
10
OpenAIGPT-6 SolNEW
80.7%
1.9K$0.00658.15s
11
80.5%
2.8K$0.00677.07s
12
80.5%
2.9K$0.00727.81s
13
79.8%
3.0K$0.007523.14s
14
79.0%
2.2K$0.008810.32s
15
78.7%
2.2K$0.0348.71s
16
AnthropicClaude Opus 5
78.3%
2.1K$0.0177.38s
17
74.9%
1.7K$0.00214.10s
18
OpenAIGPT-5.5
74.8%
2.1K$0.0229.03s
19
74.7%
2.8K~$0.000917.99s
20
73.8%
2.1K$0.00107.38s
21
73.8%
2.1K$0.00887.74s
22
73.6%
4.1K~$0.002142.09s
23
71.9%
4.6K~$0.001227.10s
24
GrokGrok 4.7NEW
71.9%
4.2K$0.01223.55s
25
70.8%
5.0K~$0.004380.37s
26
70.8%
2.3K$0.00118.70s
27
70.3%
1.6K$0.00142.70s
28
69.4%
5.6K~$0.001631.88s
29
68.8%
1.6K$0.00046.48s
30
68.7%
2.8K$0.009717.55s
31
68.7%
2.1K$0.0165.20s
32
OpenAIGPT-6 LunaNEW
68.6%
2.1K$0.000411.27s
33
67.4%
1.5K$0.00087.01s
34
67.0%
1.9K~$0.001228.79s
35
MoonshotAIKimi K3
66.5%
2.6K$0.01112.71s
36
66.4%
2.1K$0.00644.84s
37
66.3%
2.4K$0.00056.78s
38
66.0%
796$0.00506.11s
39
65.9%
1.9K$0.00079.17s
40
65.8%
2.6K$0.008420.75s
41
65.3%
2.2K$0.00316.35s
42
64.7%
1.9K$0.00305.25s
43
64.4%
1.7K$0.00077.38s
44
64.3%
5.9K~$0.001733.65s
45
64.0%
2.2K38.37s
46
63.6%
3.2K~$0.001927.84s
47
MoonshotAIKimi K2.6
62.9%
2.3K$0.003229.57s
48
62.7%
2.0K$0.00138.87s
49
61.5%
1.5K$0.00016.28s
50
61.2%
2.2K$0.00177.33s
51
59.7%
2.9K~$0.000511.49s
52
58.2%
6.8K~$0.001834.78s
53
DeepSeekDeepSeek V4 Flash Vision Exp
57.4%
703$0.00063.69s
54
47.9%
7.7K~$0.001430.60s
55
47.6%
955~$0.00024.85s
56
43.9%
969~$0.00023.60s
57
40.9%
1.7K~$0.00048.65s

Score vs. cost

Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

56 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort

Results by task

Score key:≥75%40–74%<40%
ModelObject DetectionCountingIdentificationOCRData ExtractionReasoning
82.180.289.691.988.787.2
70.680.699.089.394.582.1
AnthropicClaude Opus 5.5NEW
74.480.693.887.893.583.0
70.578.496.988.296.280.8
68.178.897.987.397.381.2
76.781.188.593.387.675.9
67.471.6100.092.695.972.2
57.180.299.088.295.977.7
61.469.497.994.093.172.0
OpenAIGPT-6 SolNEW
73.674.891.791.780.472.2
60.676.691.792.687.674.2
59.076.689.693.689.075.1
58.674.392.791.388.773.3
68.474.389.690.784.966.0
56.463.5100.094.091.866.2
AnthropicClaude Opus 5
54.470.390.693.289.771.5
38.667.693.887.696.964.9
OpenAIGPT-5.5
43.668.089.691.285.970.6
65.764.985.492.278.062.0
61.067.183.390.780.460.5
60.665.886.589.479.760.9
59.767.182.388.584.559.2
57.065.382.387.784.554.8
GrokGrok 4.7NEW
40.461.787.592.684.964.2
50.567.680.284.783.858.1
41.066.281.392.186.657.6
57.552.784.487.491.848.3
52.962.680.283.083.554.1
59.856.388.588.084.935.1
23.865.884.491.885.661.1
38.654.084.493.888.753.0
OpenAIGPT-6 LunaNEW
56.865.881.387.968.052.1
60.150.084.486.583.539.7
48.251.480.290.880.450.8
MoonshotAIKimi K3
51.946.081.393.084.542.4
36.156.881.391.789.743.0
33.155.484.490.683.551.0
33.752.793.888.884.542.4
52.147.390.688.187.629.8
19.059.583.392.183.557.6
56.548.684.489.381.431.8
15.858.683.389.584.257.0
58.854.078.184.579.431.8
46.651.883.377.578.747.7
54.351.481.381.181.634.4
44.243.281.388.776.647.7
MoonshotAIKimi K2.6
35.756.878.188.579.439.1
53.944.678.182.885.631.1
42.846.084.484.177.334.4
54.541.978.181.479.431.8
35.042.876.086.776.041.5
38.139.277.180.276.038.6
DeepSeekDeepSeek V4 Flash Vision Exp
45.346.068.886.766.031.8
31.528.863.566.769.427.1
23.822.154.283.169.832.7
19.014.453.180.463.233.1
0.033.871.951.257.131.1

Cell color follows the absolute score, on the same scale as the score key above. The column leader is bold. Hover a cell for the full breakdown.

What are Vision Evals?

Vision Evals measures what vision language models can actually do with real images. Every model receives the same samples across six tasks: Object Detection, Counting, Identification, OCR, Data Extraction, and Reasoning. Answers are scored against ground truth, so the scores are directly comparable. No human votes, no subjective judgment.

Methodology

Every model runs the full sample set for every task. Under the current benchmark protocol each task is run at two reasoning effort tiers, low and high, and each of those configurations is run three times; the reported score is the mean over the runs, and the ± after a score is half the range across them. Models benchmarked before the protocol ran once per task at low effort, plus once at high effort for Reasoning, and are being re-run under it; until then their scores are single runs and carry no ±. Every table on this site other than a task leaderboard, including the Average column, uses the low-effort pass so all models are compared on the same amount of compute. Low and high are benchmark-wide tiers, not provider settings: each model's tier is mapped to its native reasoning mechanism, whether that is an effort parameter (Claude, GPT, Grok), a thinking level (Gemini), or a thinking mode that low disables and high enables (Qwen, GLM, Kimi).

Object Detection is scored by mean Average Precision (mAP@50 as the headline, with mAP@75 and mAP@50:95 reported on the task page) and OCR by mean similarity to the ground-truth transcription. Counting, Identification, Data Extraction, and Reasoning are scored by a second model acting as judge, Gemini 3.5 Flash at temperature zero, which compares every answer with the ground truth and decides whether it is correct, so a right answer still counts when the model wraps it in explanation; a strict accuracy, counting only answers that match the ground truth exactly, is reported beside it on each task page, and the gap between the two is a measure of verbosity. The overall Average column is the unweighted mean of the six task scores.

Token usage is measured directly from each provider’s API response, and output counts include reasoning tokens. Estimated cost per sample multiplies that measured input and output usage by the model’s published per-1M pricing at the time of our last price sync, so prices stay current without re-running the benchmark; it is an estimate on this benchmark, not a universal model cost. Speed is wall-clock inference time per sample.

Last evaluated: September 22, 2026

Looking for the previous benchmark?

The previous version of Vision Evals is preserved with its full leaderboards. It covers many more models, including deprecated ones that can no longer be re-run. See Vision Evals (legacy)

Frequently Asked Questions

Each of the six tasks produces one headline score: mAP@50 for Object Detection, mean similarity for OCR, and LLM-judged accuracy for the rest. The Average is the unweighted mean of those six scores, so every task counts equally.

Strict accuracy counts an answer only when it matches the ground truth exactly after normalization. LLM-judged accuracy asks a second model acting as judge, Gemini 3.5 Flash at temperature zero, to compare the answer with the ground truth and decide whether it is correct, so a right answer still counts when it arrives with an explanation or in different words. The judge is the headline score because it measures whether the model got it right rather than whether it was terse; the strict score sits beside it on each task page, and the gap between the two shows how verbose a model is.

Under the current protocol each task is run three times at each effort tier, and the score shown is the mean over those runs. The ± is half the range between the lowest and highest run, so it shows how much the score moved between runs; it is a range, not a confidence interval. Hover the ± for the run count and range. A score without one comes from a single run; models benchmarked before the protocol are being re-run under it.

The previous version of Vision Evals is preserved at its own page and still covers many models that predate this benchmark, including deprecated models that can no longer be re-run. You can browse it at /evals/legacy.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.