Roboflow

Visual Reasoning Benchmark

The Reasoning task asks each model questions that require thinking beyond simply reading or spotting something. It covers arithmetic, comparisons, spatial relationships, logical deduction, and multi-step inference that combines several visual facts. The answer is never written directly in the image; the model must interpret the scene and reason about it.

30 models evaluated60 runs across effort levels

Evals updated August 14, 2026Pricing updated August 18, 2026

Score key:≥75%40–74%<40%
1low
84.1%
1.8K$0.00805.13s
2low
82.8%
1.5K$0.00118.62s
2high
82.8%
2.9K$0.0178.78s
4high
82.1%
2.3K$0.00269.24s
5high
80.8%
2.5K$0.01132.95s
6low
80.1%
1.7K$0.00314.51s
6high
80.1%
3.2K$0.00859.68s
8high
76.2%
3.7K$0.01216.31s
8high
76.2%
4.2K$0.01321.97s
10low
74.8%
2.6K$0.006510.11s
10low
74.8%
2.7K$0.007410.27s
10high
74.8%
2.7K$0.02113.32s
13high
74.2%
2.3K$0.00406.41s
13
AnthropicClaude Opus 5
high
74.2%
1.8K$0.0188.59s
13
MoonshotAIKimi K3
high
74.2%
3.6K$0.03779.92s
16low
73.5%
1.5K$0.004711.76s
17high
72.2%
1.6K$0.00798.01s
17low
72.2%
1.9K$0.0129.27s
19
AnthropicClaude Opus 5
low
71.5%
1.5K$0.0104.99s
20
OpenAIGPT-5.5
high
69.5%
2.0K$0.03015.40s
21high
68.9%
2.7K$0.00425.82s
22
QwenQwen 3.7 Plus
high
68.2%
4.1K$0.004358.41s
23
OpenAIGPT-5.5
low
67.5%
1.5K$0.0146.64s
24low
66.2%
1.5K$0.0185.80s
24high
66.2%
1.6K$0.0246.88s
26low
65.6%
1.4K$0.00575.20s
27low
64.9%
1.6K$0.00203.47s
28high
64.2%
1.6K$0.00665.38s
29high
62.9%
2.8K$0.008113.25s
29high
62.9%
3.3K$0.003325.13s
31high
62.3%
1.3K$0.0119.30s
31high
62.3%
3.7K$0.008753.85s
33
GrokGrok 4.6NEW
low
61.6%
2.3K$0.008710.29s
33
GrokGrok 4.6NEW
high
61.6%
5.3K$0.02737.83s
33high
61.6%
5.1K$0.006562.64s
36high
60.9%
2.3K$0.00158.49s
36high
60.9%
4.6K$0.000559.29s
38low
59.6%
1.5K$0.00506.43s
38high
59.6%
2.8K$0.01128.67s
40
MoonshotAIKimi K2.6
high
58.9%
6.8K$0.020111.19s
41low
58.3%
2.3K$0.007616.48s
42low
57.6%
1.8K$0.00106.55s
43low
55.6%
1.6K$0.00234.07s
44low
55.0%
1.6K$0.00064.85s
45low
53.0%
1.4K$0.00783.00s
46high
52.3%
1.4K$0.00782.45s
47
Z.aiGLM 5V Turbo
high
49.7%
2.7K$0.006929.02s
48low
48.3%
1.5K$0.00122.38s
49low
43.0%
1.4K$0.00323.09s
49high
43.0%
1.6K$0.00434.12s
51low
42.4%
403$0.00133.98s
51
MoonshotAIKimi K3
low
42.4%
1.4K$0.00444.51s
53
QwenQwen 3.7 Plus
low
39.7%
1.1K$0.00032.19s
54
MoonshotAIKimi K2.6
low
39.1%
1.5K$0.00133.38s
55low
34.4%
1.1K<$0.00013.70s
56high
33.8%
1.1K$0.00021.96s
57low
31.8%
1.1K$0.00051.08s
57low
31.8%
1.0K$0.00022.16s
57
Z.aiGLM 5V Turbo
low
31.8%
1.4K$0.00173.77s
60low
29.8%
1.1K$0.00021.69s

Score vs. cost

Reasoning score (Accuracy) against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.

30 models on the current benchmark · Reasoning task only, low effort

Example Reasoning benchmark tasks

Real samples from the benchmark: the image each model sees, the question it is asked, and the ground-truth answer it is scored against.

Benchmark sample: Pipe Reordering (Minimum Moves)

The models are asked

If a worker wants to rearrange the five numbered pipes so that their labels are in strictly ascending order from top to bottom, what is the minimum number of pipes that must be moved? Answer with a single integer.

Ground truth

1

Benchmark sample: Conditional Probability

The models are asked

Given that the selected number satisfies condition (iii), what is the probability that it also satisfies condition (iv)? Choose one: 0, 1/4, 1/2, or 1.

Ground truth

1/2

Benchmark sample: Technical Drawing (Measurement)

The models are asked

Based on the technical drawing, what is the horizontal distance between the leftmost center of the left slot and the vertical centerline of the top hole? Answer with a number rounded to three decimal places.

Ground truth

2.742

Benchmark sample: Collision Perspective (Left/right)

The models are asked

From the perspective of a driver sitting inside the white sedan and looking forward, on which side of their vehicle's front end did the collision occur? Choose one: left or right.

Ground truth

left

Benchmark sample: Appointment Scheduling (Time Slots)

The models are asked

If a patient can only attend appointments at or after 11:00 AM, how many of the displayed time slots across the entire week are available to them? Answer with a single integer.

Ground truth

10

How Reasoning is scored

Each answer is graded against the ground truth. The leaderboard score is the percentage of samples answered correctly. Reasoning is the one task run at two reasoning effort levels, so each model appears once for its low-effort pass and once for its high-effort pass, ranked together on the same samples.

Every model runs the same sample set in a single evaluation pass. Token usage is measured from each provider’s API response, and cost per sample is that usage multiplied by the model’s published pricing. See the full methodology.

Frequently Asked Questions

Each model answers the same reasoning prompts. Answers are graded against the ground truth, and the score is the percentage answered correctly.

Reasoning is run at two levels of reasoning effort. The low row is the model answering with minimal deliberation, the high row is the same model on the same questions allowed to think longer. Both rows are ranked together so you can see whether the extra thinking is worth its cost and latency, which the Est. cost and Speed columns show. For most models it helps, for a few it changes nothing at all, and for a couple the score drops.

Low and high are benchmark tiers mapped to each model's native reasoning mechanism, because no two labs expose the same control. Claude, GPT, and Grok take low or high as a reasoning effort parameter. Gemini runs at its low and high thinking levels (a minimal versus large thinking budget for the older Gemini 2.5 Pro). Qwen, GLM, and Kimi only expose an on/off thinking mode, so low runs them with thinking disabled and high with thinking at high effort. Muse and Qwen3.8-Max cannot switch reasoning off, so their low pass uses the lowest native setting. All six tasks and the Average use the low pass; high has been run only for Reasoning.

Every other table, including the overall Average, uses the low-effort pass. High effort has only been run for Reasoning, so using it elsewhere would compare models on different amounts of compute.

The answer is never written directly in the image. The model has to combine several visual facts, such as reading two values and comparing them, judging spatial relationships, or applying logic to what it sees.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.