Visual Reasoning Benchmark
The Reasoning task asks each model questions that require thinking beyond simply reading or spotting something. It covers arithmetic, comparisons, spatial relationships, logical deduction, and multi-step inference that combines several visual facts. The answer is never written directly in the image; the model must interpret the scene and reason about it.
Evals updated August 14, 2026Pricing updated August 18, 2026
| 1 | low | 84.1% | 1.8K | $0.0080 | 5.13s | ||
| 2 | low | 82.8% | 1.5K | $0.0011 | 8.62s | ||
| 2 | high | 82.8% | 2.9K | $0.017 | 8.78s | ||
| 4 | high | 82.1% | 2.3K | $0.0026 | 9.24s | ||
| 5 | high | 80.8% | 2.5K | $0.011 | 32.95s | ||
| 6 | low | 80.1% | 1.7K | $0.0031 | 4.51s | ||
| 6 | high | 80.1% | 3.2K | $0.0085 | 9.68s | ||
| 8 | high | 76.2% | 3.7K | $0.012 | 16.31s | ||
| 8 | high | 76.2% | 4.2K | $0.013 | 21.97s | ||
| 10 | low | 74.8% | 2.6K | $0.0065 | 10.11s | ||
| 10 | low | 74.8% | 2.7K | $0.0074 | 10.27s | ||
| 10 | high | 74.8% | 2.7K | $0.021 | 13.32s | ||
| 13 | high | 74.2% | 2.3K | $0.0040 | 6.41s | ||
| 13 | high | 74.2% | 1.8K | $0.018 | 8.59s | ||
| 13 | high | 74.2% | 3.6K | $0.037 | 79.92s | ||
| 16 | low | 73.5% | 1.5K | $0.0047 | 11.76s | ||
| 17 | high | 72.2% | 1.6K | $0.0079 | 8.01s | ||
| 17 | low | 72.2% | 1.9K | $0.012 | 9.27s | ||
| 19 | low | 71.5% | 1.5K | $0.010 | 4.99s | ||
| 20 | high | 69.5% | 2.0K | $0.030 | 15.40s | ||
| 21 | high | 68.9% | 2.7K | $0.0042 | 5.82s | ||
| 22 | Qwen 3.7 Plus | high | 68.2% | 4.1K | $0.0043 | 58.41s | |
| 23 | low | 67.5% | 1.5K | $0.014 | 6.64s | ||
| 24 | low | 66.2% | 1.5K | $0.018 | 5.80s | ||
| 24 | high | 66.2% | 1.6K | $0.024 | 6.88s | ||
| 26 | low | 65.6% | 1.4K | $0.0057 | 5.20s | ||
| 27 | low | 64.9% | 1.6K | $0.0020 | 3.47s | ||
| 28 | high | 64.2% | 1.6K | $0.0066 | 5.38s | ||
| 29 | high | 62.9% | 2.8K | $0.0081 | 13.25s | ||
| 29 | high | 62.9% | 3.3K | $0.0033 | 25.13s | ||
| 31 | high | 62.3% | 1.3K | $0.011 | 9.30s | ||
| 31 | Qwen3.8 27BNEW | high | 62.3% | 3.7K | $0.0087 | 53.85s | |
| 33 | Grok 4.6NEW | low | 61.6% | 2.3K | $0.0087 | 10.29s | |
| 33 | Grok 4.6NEW | high | 61.6% | 5.3K | $0.027 | 37.83s | |
| 33 | high | 61.6% | 5.1K | $0.0065 | 62.64s | ||
| 36 | high | 60.9% | 2.3K | $0.0015 | 8.49s | ||
| 36 | high | 60.9% | 4.6K | $0.0005 | 59.29s | ||
| 38 | low | 59.6% | 1.5K | $0.0050 | 6.43s | ||
| 38 | high | 59.6% | 2.8K | $0.011 | 28.67s | ||
| 40 | Kimi K2.6 | high | 58.9% | 6.8K | $0.020 | 111.19s | |
| 41 | low | 58.3% | 2.3K | $0.0076 | 16.48s | ||
| 42 | low | 57.6% | 1.8K | $0.0010 | 6.55s | ||
| 43 | low | 55.6% | 1.6K | $0.0023 | 4.07s | ||
| 44 | low | 55.0% | 1.6K | $0.0006 | 4.85s | ||
| 45 | low | 53.0% | 1.4K | $0.0078 | 3.00s | ||
| 46 | high | 52.3% | 1.4K | $0.0078 | 2.45s | ||
| 47 | GLM 5V Turbo | high | 49.7% | 2.7K | $0.0069 | 29.02s | |
| 48 | low | 48.3% | 1.5K | $0.0012 | 2.38s | ||
| 49 | low | 43.0% | 1.4K | $0.0032 | 3.09s | ||
| 49 | high | 43.0% | 1.6K | $0.0043 | 4.12s | ||
| 51 | low | 42.4% | 403 | $0.0013 | 3.98s | ||
| 51 | low | 42.4% | 1.4K | $0.0044 | 4.51s | ||
| 53 | Qwen 3.7 Plus | low | 39.7% | 1.1K | $0.0003 | 2.19s | |
| 54 | Kimi K2.6 | low | 39.1% | 1.5K | $0.0013 | 3.38s | |
| 55 | low | 34.4% | 1.1K | <$0.0001 | 3.70s | ||
| 56 | high | 33.8% | 1.1K | $0.0002 | 1.96s | ||
| 57 | Qwen3.8 27BNEW | low | 31.8% | 1.1K | $0.0005 | 1.08s | |
| 57 | low | 31.8% | 1.0K | $0.0002 | 2.16s | ||
| 57 | GLM 5V Turbo | low | 31.8% | 1.4K | $0.0017 | 3.77s | |
| 60 | low | 29.8% | 1.1K | $0.0002 | 1.69s |
Score vs. cost
Reasoning score (Accuracy) against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
30 models on the current benchmark · Reasoning task only, low effort
Example Reasoning benchmark tasks
Real samples from the benchmark: the image each model sees, the question it is asked, and the ground-truth answer it is scored against.






The models are asked
If a worker wants to rearrange the five numbered pipes so that their labels are in strictly ascending order from top to bottom, what is the minimum number of pipes that must be moved? Answer with a single integer.
Ground truth
1

The models are asked
Given that the selected number satisfies condition (iii), what is the probability that it also satisfies condition (iv)? Choose one: 0, 1/4, 1/2, or 1.
Ground truth
1/2

The models are asked
Based on the technical drawing, what is the horizontal distance between the leftmost center of the left slot and the vertical centerline of the top hole? Answer with a number rounded to three decimal places.
Ground truth
2.742

The models are asked
From the perspective of a driver sitting inside the white sedan and looking forward, on which side of their vehicle's front end did the collision occur? Choose one: left or right.
Ground truth
left

The models are asked
If a patient can only attend appointments at or after 11:00 AM, how many of the displayed time slots across the entire week are available to them? Answer with a single integer.
Ground truth
10
How Reasoning is scored
Each answer is graded against the ground truth. The leaderboard score is the percentage of samples answered correctly. Reasoning is the one task run at two reasoning effort levels, so each model appears once for its low-effort pass and once for its high-effort pass, ranked together on the same samples.
Every model runs the same sample set in a single evaluation pass. Token usage is measured from each provider’s API response, and cost per sample is that usage multiplied by the model’s published pricing. See the full methodology.
Frequently Asked Questions
Each model answers the same reasoning prompts. Answers are graded against the ground truth, and the score is the percentage answered correctly.
Reasoning is run at two levels of reasoning effort. The low row is the model answering with minimal deliberation, the high row is the same model on the same questions allowed to think longer. Both rows are ranked together so you can see whether the extra thinking is worth its cost and latency, which the Est. cost and Speed columns show. For most models it helps, for a few it changes nothing at all, and for a couple the score drops.
Low and high are benchmark tiers mapped to each model's native reasoning mechanism, because no two labs expose the same control. Claude, GPT, and Grok take low or high as a reasoning effort parameter. Gemini runs at its low and high thinking levels (a minimal versus large thinking budget for the older Gemini 2.5 Pro). Qwen, GLM, and Kimi only expose an on/off thinking mode, so low runs them with thinking disabled and high with thinking at high effort. Muse and Qwen3.8-Max cannot switch reasoning off, so their low pass uses the lowest native setting. All six tasks and the Average use the low pass; high has been run only for Reasoning.
Every other table, including the overall Average, uses the low-effort pass. High effort has only been run for Reasoning, so using it elsewhere would compare models on different amounts of compute.
The answer is never written directly in the image. The model has to combine several visual facts, such as reading two values and comparing them, judging spatial relationships, or applying logic to what it sees.
Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.