Vision Evals
See which AI vision models are best at object detection, counting, identification, OCR, data extraction, and reasoning. Every model runs the same real-world tasks and is scored against ground truth by Roboflow.
Evals updated September 22, 2026Pricing updated September 22, 2026
Overall ranking
| 1 | 86.6% | 1.9K | $0.030 | 6.67s | ||
| 2 | 86.0% | 2.1K | $0.011 | 14.77s | ||
| 3 | 85.5% | 2.2K | $0.014 | 12.76s | ||
| 4 | 85.2% | 1.8K | $0.0031 | 16.50s | ||
| 5 | 85.1% | 1.8K | $0.0033 | 11.65s | ||
| 6 | 83.9% | 2.1K | $0.0074 | 17.25s | ||
| 7 | 83.3% | 1.7K | $0.0093 | 7.81s | ||
| 8 | 83.0% | 1.8K | $0.0032 | 14.66s | ||
| 9 | 81.3% | 2.2K | $0.035 | 8.28s | ||
| 10 | GPT-6 SolNEW | 80.7% | 1.9K | $0.0065 | 8.15s | |
| 11 | 80.5% | 2.8K | $0.0067 | 7.07s | ||
| 12 | 80.5% | 2.9K | $0.0072 | 7.81s | ||
| 13 | 79.8% | 3.0K | $0.0075 | 23.14s | ||
| 14 | 79.0% | 2.2K | $0.0088 | 10.32s | ||
| 15 | 78.7% | 2.2K | $0.034 | 8.71s | ||
| 16 | 78.3% | 2.1K | $0.017 | 7.38s | ||
| 17 | 74.9% | 1.7K | $0.0021 | 4.10s | ||
| 18 | 74.8% | 2.1K | $0.022 | 9.03s | ||
| 19 | 74.7% | 2.8K | ~$0.0009 | 17.99s | ||
| 20 | 73.8% | 2.1K | $0.0010 | 7.38s | ||
| 21 | 73.8% | 2.1K | $0.0088 | 7.74s | ||
| 22 | 73.6% | 4.1K | ~$0.0021 | 42.09s | ||
| 23 | 71.9% | 4.6K | ~$0.0012 | 27.10s | ||
| 24 | Grok 4.7NEW | 71.9% | 4.2K | $0.012 | 23.55s | |
| 25 | 70.8% | 5.0K | ~$0.0043 | 80.37s | ||
| 26 | 70.8% | 2.3K | $0.0011 | 8.70s | ||
| 27 | 70.3% | 1.6K | $0.0014 | 2.70s | ||
| 28 | 69.4% | 5.6K | ~$0.0016 | 31.88s | ||
| 29 | 68.8% | 1.6K | $0.0004 | 6.48s | ||
| 30 | 68.7% | 2.8K | $0.0097 | 17.55s | ||
| 31 | 68.7% | 2.1K | $0.016 | 5.20s | ||
| 32 | GPT-6 LunaNEW | 68.6% | 2.1K | $0.0004 | 11.27s | |
| 33 | 67.4% | 1.5K | $0.0008 | 7.01s | ||
| 34 | 67.0% | 1.9K | ~$0.0012 | 28.79s | ||
| 35 | 66.5% | 2.6K | $0.011 | 12.71s | ||
| 36 | 66.4% | 2.1K | $0.0064 | 4.84s | ||
| 37 | 66.3% | 2.4K | $0.0005 | 6.78s | ||
| 38 | 66.0% | 796 | $0.0050 | 6.11s | ||
| 39 | 65.9% | 1.9K | $0.0007 | 9.17s | ||
| 40 | 65.8% | 2.6K | $0.0084 | 20.75s | ||
| 41 | 65.3% | 2.2K | $0.0031 | 6.35s | ||
| 42 | 64.7% | 1.9K | $0.0030 | 5.25s | ||
| 43 | 64.4% | 1.7K | $0.0007 | 7.38s | ||
| 44 | 64.3% | 5.9K | ~$0.0017 | 33.65s | ||
| 45 | 64.0% | 2.2K | — | 38.37s | ||
| 46 | 63.6% | 3.2K | ~$0.0019 | 27.84s | ||
| 47 | Kimi K2.6 | 62.9% | 2.3K | $0.0032 | 29.57s | |
| 48 | 62.7% | 2.0K | $0.0013 | 8.87s | ||
| 49 | 61.5% | 1.5K | $0.0001 | 6.28s | ||
| 50 | 61.2% | 2.2K | $0.0017 | 7.33s | ||
| 51 | 59.7% | 2.9K | ~$0.0005 | 11.49s | ||
| 52 | 58.2% | 6.8K | ~$0.0018 | 34.78s | ||
| 53 | DeepSeek V4 Flash Vision Exp | 57.4% | 703 | $0.0006 | 3.69s | |
| 54 | 47.9% | 7.7K | ~$0.0014 | 30.60s | ||
| 55 | 47.6% | 955 | ~$0.0002 | 4.85s | ||
| 56 | 43.9% | 969 | ~$0.0002 | 3.60s | ||
| 57 | 40.9% | 1.7K | ~$0.0004 | 8.65s |
Score vs. cost
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
56 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort
Results by task
Object Detection
Locate and classify objects with a bounding box and label for each one.
- 1GPT-6 Astra82.1%
- 2Qwen3.8 Max76.7%
- 3Claude Opus 5.574.4%
Counting
Report exactly how many of a specified thing appear in an image.
- 1Qwen3.8 Max81.1%
- 2Claude Opus 5.580.6%
- 2Gemini 3.5 Flash80.6%
Identification
Recognize and name a specific entity in an image from visual evidence.
- 1Claude Fable 5100.0%
- 1Gemini 3.1 Pro100.0%
- 3Gemini 3.5 Flash99.0%
OCR
Read and transcribe every piece of text in an image, exactly as it appears.
- 1Claude Fable 594.0%
- 2Claude Fable 5.194.0%
- 3Claude Opus 4.893.8%
Data Extraction
Find and return one specific piece of information from an image.
- 1Gemini 3.8 Flash97.3%
- 2Gemini 3 Flash96.9%
- 3Gemini 3.7 Flash96.2%
Reasoning
Answer questions that require thinking beyond reading or spotting something.
- 1GPT-6 Astra87.2%
- 2Claude Opus 5.583.0%
- 3Gemini 3.5 Flash82.1%
| Model | Object Detection | Counting | Identification | OCR | Data Extraction | Reasoning |
|---|---|---|---|---|---|---|
| 82.1 | 80.2 | 89.6 | 91.9 | 88.7 | 87.2 | |
| 70.6 | 80.6 | 99.0 | 89.3 | 94.5 | 82.1 | |
| 74.4 | 80.6 | 93.8 | 87.8 | 93.5 | 83.0 | |
| 70.5 | 78.4 | 96.9 | 88.2 | 96.2 | 80.8 | |
| 68.1 | 78.8 | 97.9 | 87.3 | 97.3 | 81.2 | |
| 76.7 | 81.1 | 88.5 | 93.3 | 87.6 | 75.9 | |
| 67.4 | 71.6 | 100.0 | 92.6 | 95.9 | 72.2 | |
| 57.1 | 80.2 | 99.0 | 88.2 | 95.9 | 77.7 | |
| 61.4 | 69.4 | 97.9 | 94.0 | 93.1 | 72.0 | |
GPT-6 SolNEW | 73.6 | 74.8 | 91.7 | 91.7 | 80.4 | 72.2 |
| 60.6 | 76.6 | 91.7 | 92.6 | 87.6 | 74.2 | |
| 59.0 | 76.6 | 89.6 | 93.6 | 89.0 | 75.1 | |
| 58.6 | 74.3 | 92.7 | 91.3 | 88.7 | 73.3 | |
| 68.4 | 74.3 | 89.6 | 90.7 | 84.9 | 66.0 | |
| 56.4 | 63.5 | 100.0 | 94.0 | 91.8 | 66.2 | |
| 54.4 | 70.3 | 90.6 | 93.2 | 89.7 | 71.5 | |
| 38.6 | 67.6 | 93.8 | 87.6 | 96.9 | 64.9 | |
| 43.6 | 68.0 | 89.6 | 91.2 | 85.9 | 70.6 | |
| 65.7 | 64.9 | 85.4 | 92.2 | 78.0 | 62.0 | |
| 61.0 | 67.1 | 83.3 | 90.7 | 80.4 | 60.5 | |
| 60.6 | 65.8 | 86.5 | 89.4 | 79.7 | 60.9 | |
| 59.7 | 67.1 | 82.3 | 88.5 | 84.5 | 59.2 | |
| 57.0 | 65.3 | 82.3 | 87.7 | 84.5 | 54.8 | |
Grok 4.7NEW | 40.4 | 61.7 | 87.5 | 92.6 | 84.9 | 64.2 |
| 50.5 | 67.6 | 80.2 | 84.7 | 83.8 | 58.1 | |
| 41.0 | 66.2 | 81.3 | 92.1 | 86.6 | 57.6 | |
| 57.5 | 52.7 | 84.4 | 87.4 | 91.8 | 48.3 | |
| 52.9 | 62.6 | 80.2 | 83.0 | 83.5 | 54.1 | |
| 59.8 | 56.3 | 88.5 | 88.0 | 84.9 | 35.1 | |
| 23.8 | 65.8 | 84.4 | 91.8 | 85.6 | 61.1 | |
| 38.6 | 54.0 | 84.4 | 93.8 | 88.7 | 53.0 | |
GPT-6 LunaNEW | 56.8 | 65.8 | 81.3 | 87.9 | 68.0 | 52.1 |
| 60.1 | 50.0 | 84.4 | 86.5 | 83.5 | 39.7 | |
| 48.2 | 51.4 | 80.2 | 90.8 | 80.4 | 50.8 | |
| 51.9 | 46.0 | 81.3 | 93.0 | 84.5 | 42.4 | |
| 36.1 | 56.8 | 81.3 | 91.7 | 89.7 | 43.0 | |
| 33.1 | 55.4 | 84.4 | 90.6 | 83.5 | 51.0 | |
| 33.7 | 52.7 | 93.8 | 88.8 | 84.5 | 42.4 | |
| 52.1 | 47.3 | 90.6 | 88.1 | 87.6 | 29.8 | |
| 19.0 | 59.5 | 83.3 | 92.1 | 83.5 | 57.6 | |
| 56.5 | 48.6 | 84.4 | 89.3 | 81.4 | 31.8 | |
| 15.8 | 58.6 | 83.3 | 89.5 | 84.2 | 57.0 | |
| 58.8 | 54.0 | 78.1 | 84.5 | 79.4 | 31.8 | |
| 46.6 | 51.8 | 83.3 | 77.5 | 78.7 | 47.7 | |
| 54.3 | 51.4 | 81.3 | 81.1 | 81.6 | 34.4 | |
| 44.2 | 43.2 | 81.3 | 88.7 | 76.6 | 47.7 | |
Kimi K2.6 | 35.7 | 56.8 | 78.1 | 88.5 | 79.4 | 39.1 |
| 53.9 | 44.6 | 78.1 | 82.8 | 85.6 | 31.1 | |
| 42.8 | 46.0 | 84.4 | 84.1 | 77.3 | 34.4 | |
| 54.5 | 41.9 | 78.1 | 81.4 | 79.4 | 31.8 | |
| 35.0 | 42.8 | 76.0 | 86.7 | 76.0 | 41.5 | |
| 38.1 | 39.2 | 77.1 | 80.2 | 76.0 | 38.6 | |
DeepSeek V4 Flash Vision Exp | 45.3 | 46.0 | 68.8 | 86.7 | 66.0 | 31.8 |
| 31.5 | 28.8 | 63.5 | 66.7 | 69.4 | 27.1 | |
| 23.8 | 22.1 | 54.2 | 83.1 | 69.8 | 32.7 | |
| 19.0 | 14.4 | 53.1 | 80.4 | 63.2 | 33.1 | |
| 0.0 | 33.8 | 71.9 | 51.2 | 57.1 | 31.1 |
Cell color follows the absolute score, on the same scale as the score key above. The column leader is bold. Hover a cell for the full breakdown.
What are Vision Evals?
Vision Evals measures what vision language models can actually do with real images. Every model receives the same samples across six tasks: Object Detection, Counting, Identification, OCR, Data Extraction, and Reasoning. Answers are scored against ground truth, so the scores are directly comparable. No human votes, no subjective judgment.
Methodology
Every model runs the full sample set for every task. Under the current benchmark protocol each task is run at two reasoning effort tiers, low and high, and each of those configurations is run three times; the reported score is the mean over the runs, and the ± after a score is half the range across them. Models benchmarked before the protocol ran once per task at low effort, plus once at high effort for Reasoning, and are being re-run under it; until then their scores are single runs and carry no ±. Every table on this site other than a task leaderboard, including the Average column, uses the low-effort pass so all models are compared on the same amount of compute. Low and high are benchmark-wide tiers, not provider settings: each model's tier is mapped to its native reasoning mechanism, whether that is an effort parameter (Claude, GPT, Grok), a thinking level (Gemini), or a thinking mode that low disables and high enables (Qwen, GLM, Kimi).
Object Detection is scored by mean Average Precision (mAP@50 as the headline, with mAP@75 and mAP@50:95 reported on the task page) and OCR by mean similarity to the ground-truth transcription. Counting, Identification, Data Extraction, and Reasoning are scored by a second model acting as judge, Gemini 3.5 Flash at temperature zero, which compares every answer with the ground truth and decides whether it is correct, so a right answer still counts when the model wraps it in explanation; a strict accuracy, counting only answers that match the ground truth exactly, is reported beside it on each task page, and the gap between the two is a measure of verbosity. The overall Average column is the unweighted mean of the six task scores.
Token usage is measured directly from each provider’s API response, and output counts include reasoning tokens. Estimated cost per sample multiplies that measured input and output usage by the model’s published per-1M pricing at the time of our last price sync, so prices stay current without re-running the benchmark; it is an estimate on this benchmark, not a universal model cost. Speed is wall-clock inference time per sample.
Last evaluated: September 22, 2026
Looking for the previous benchmark?
The previous version of Vision Evals is preserved with its full leaderboards. It covers many more models, including deprecated ones that can no longer be re-run. See Vision Evals (legacy)
Frequently Asked Questions
Each of the six tasks produces one headline score: mAP@50 for Object Detection, mean similarity for OCR, and LLM-judged accuracy for the rest. The Average is the unweighted mean of those six scores, so every task counts equally.
Strict accuracy counts an answer only when it matches the ground truth exactly after normalization. LLM-judged accuracy asks a second model acting as judge, Gemini 3.5 Flash at temperature zero, to compare the answer with the ground truth and decide whether it is correct, so a right answer still counts when it arrives with an explanation or in different words. The judge is the headline score because it measures whether the model got it right rather than whether it was terse; the strict score sits beside it on each task page, and the gap between the two shows how verbose a model is.
Under the current protocol each task is run three times at each effort tier, and the score shown is the mean over those runs. The ± is half the range between the lowest and highest run, so it shows how much the score moved between runs; it is a range, not a confidence interval. Hover the ± for the run count and range. A score without one comes from a single run; models benchmarked before the protocol are being re-run under it.
The previous version of Vision Evals is preserved at its own page and still covers many models that predate this benchmark, including deprecated models that can no longer be re-run. You can browse it at /evals/legacy.
Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.