Roboflow

Object Detection Benchmark

The Object Detection task asks each model to locate and classify objects by returning a bounding box and a class label for every one, in the requested output format. Unlike counting or identification, the model must report both what each object is and precisely where it sits in the image. Each prompt restricts the model to a fixed list of class labels.

53 models evaluated70 runs across effort levels

Evals updated September 5, 2026Pricing updated September 13, 2026

Score key:≥75%40–74%<40%
1
OpenAIGPT-6 AstraNEW
high
83.6%
±0.8, Mean of 3 runs, range 82.8 to 84.5
67.0%±0.5, Mean of 3 runs, range 66.5 to 67.561.9%±0.2, Mean of 3 runs, range 61.7 to 62.13.6K$0.10131.66s
2
OpenAIGPT-6 AstraNEW
low
82.1%
±0.8, Mean of 3 runs, range 81.0 to 82.7
63.8%±0.9, Mean of 3 runs, range 62.9 to 64.859.1%±1.1, Mean of 3 runs, range 57.8 to 60.02.6K$0.05010.96s
3high
78.4%
±0.4, Mean of 3 runs, range 78.1 to 78.9
64.6%±0.9, Mean of 3 runs, range 63.9 to 65.862.3%±0.5, Mean of 3 runs, range 61.9 to 63.06.0K$0.03083.04s
4low
76.7%
±0.3, Mean of 3 runs, range 76.5 to 77.1
61.4%±0.3, Mean of 3 runs, range 61.2 to 61.759.9%±0.3, Mean of 3 runs, range 59.6 to 60.33.1K$0.01229.91s
5high
74.8%
±1.1, Mean of 3 runs, range 73.4 to 75.6
62.5%±0.4, Mean of 3 runs, range 62.0 to 62.859.0%±0.7, Mean of 3 runs, range 58.1 to 59.56.6K$0.02120.96s
6high
74.3%
±0.8, Mean of 3 runs, range 73.3 to 75.0
62.1%±1.3, Mean of 3 runs, range 60.9 to 63.558.9%±0.9, Mean of 3 runs, range 58.1 to 59.93.3K$0.008922.34s
7high
70.7%
±0.4, Mean of 3 runs, range 70.3 to 71.2
58.1%±0.7, Mean of 3 runs, range 57.6 to 59.055.3%±0.5, Mean of 3 runs, range 54.9 to 56.03.4K$0.009326.16s
8low
70.6%
±2.0, Mean of 3 runs, range 68.7 to 72.6
56.8%±2.0, Mean of 3 runs, range 55.4 to 59.354.4%±1.7, Mean of 3 runs, range 53.0 to 56.42.7K$0.01620.60s
9low
70.5%
±1.1, Mean of 3 runs, range 69.4 to 71.5
59.3%±0.7, Mean of 3 runs, range 58.5 to 59.856.4%±0.5, Mean of 3 runs, range 56.0 to 57.02.2K$0.004719.71s
10high
69.8%
±1.8, Mean of 3 runs, range 67.5 to 71.1
57.1%±2.1, Mean of 3 runs, range 54.6 to 58.954.3%±1.8, Mean of 3 runs, range 52.0 to 55.63.3K$0.02123.88s
11low
68.4%
±0.7, Mean of 3 runs, range 67.9 to 69.3
46.1%±1.4, Mean of 3 runs, range 44.5 to 47.345.2%±0.5, Mean of 3 runs, range 44.7 to 45.73.2K$0.01518.11s
12high
68.4%
±0.8, Mean of 3 runs, range 67.7 to 69.3
45.1%±0.5, Mean of 3 runs, range 44.5 to 45.444.6%±0.5, Mean of 3 runs, range 44.2 to 45.35.1K$0.03543.12s
13low
68.1%
±0.8, Mean of 3 runs, range 67.3 to 69.0
57.3%±0.5, Mean of 3 runs, range 56.7 to 57.654.5%±0.5, Mean of 3 runs, range 54.0 to 55.02.0K$0.004117.01s
14low
67.4%
55.8%52.4%1.9K$0.0108.76s
15high
67.0%
±1.6, Mean of 3 runs, range 65.3 to 68.5
53.8%±0.8, Mean of 3 runs, range 52.9 to 54.551.9%±1.0, Mean of 3 runs, range 50.8 to 52.83.3K$0.001015.78s
16high
66.1%
±1.4, Mean of 3 runs, range 64.9 to 67.8
50.0%±0.9, Mean of 3 runs, range 49.3 to 51.149.1%±0.9, Mean of 3 runs, range 48.5 to 50.25.2K$039.37s
17low
65.7%
±1.0, Mean of 3 runs, range 64.6 to 66.5
49.8%±0.9, Mean of 3 runs, range 49.0 to 50.948.8%±0.7, Mean of 3 runs, range 48.3 to 49.74.8K$033.43s
18
AnthropicClaude Fable 5.1NEW
high
65.0%
±0.4, Mean of 3 runs, range 64.6 to 65.3
40.4%±1.1, Mean of 3 runs, range 39.5 to 41.739.6%±0.5, Mean of 3 runs, range 39.2 to 40.13.5K$0.07817.13s
19high
62.3%
±1.2, Mean of 3 runs, range 61.4 to 63.8
43.4%±1.0, Mean of 3 runs, range 42.8 to 44.742.6%±1.1, Mean of 3 runs, range 41.9 to 44.05.8K$0.005038.49s
20
AnthropicClaude Fable 5.1NEW
low
61.4%
±0.5, Mean of 3 runs, range 61.0 to 62.0
40.2%±0.9, Mean of 3 runs, range 39.3 to 41.038.0%±0.9, Mean of 3 runs, range 37.3 to 39.13.2K$0.06011.98s
21high
61.3%
±0.4, Mean of 3 runs, range 60.8 to 61.6
39.1%±1.1, Mean of 3 runs, range 38.3 to 40.639.0%±0.2, Mean of 3 runs, range 38.9 to 39.23.9K$0.02625.23s
22low
61.0%
±1.2, Mean of 3 runs, range 59.9 to 62.2
42.5%±1.5, Mean of 3 runs, range 41.4 to 44.442.0%±1.2, Mean of 3 runs, range 41.1 to 43.53.0K$0.001511.19s
23low
60.6%
±1.9, Mean of 3 runs, range 58.4 to 62.1
40.7%±1.4, Mean of 3 runs, range 39.4 to 42.239.4%±0.9, Mean of 3 runs, range 38.6 to 40.43.9K$0.00979.44s
24low
60.6%
±0.2, Mean of 3 runs, range 60.3 to 60.7
40.6%±0.3, Mean of 3 runs, range 40.4 to 40.939.6%±0.0, Mean of 3 runs, range 39.6 to 39.62.9K$0.01411.33s
25high
60.5%
±0.3, Mean of 3 runs, range 60.2 to 60.7
42.1%±1.9, Mean of 3 runs, range 40.5 to 44.239.5%±0.8, Mean of 3 runs, range 38.7 to 40.34.8K$0.01415.48s
26low
60.1%
48.4%45.1%2.2K$0.001313.21s
27high
60.0%
±0.5, Mean of 3 runs, range 59.6 to 60.7
40.9%±0.7, Mean of 3 runs, range 40.0 to 41.538.5%±0.1, Mean of 3 runs, range 38.5 to 38.65.1K$0.01512.00s
28low
59.8%
±1.1, Mean of 3 runs, range 58.5 to 60.8
48.2%±1.5, Mean of 3 runs, range 47.0 to 50.046.4%±1.3, Mean of 3 runs, range 45.0 to 47.72.4K$0.000610.37s
29low
59.7%
±0.9, Mean of 3 runs, range 59.0 to 60.8
48.5%±0.5, Mean of 3 runs, range 47.8 to 48.946.2%±0.3, Mean of 3 runs, range 45.9 to 46.53.8K$017.47s
30low
59.0%
±1.0, Mean of 3 runs, range 58.1 to 60.2
41.6%±1.7, Mean of 3 runs, range 40.1 to 43.539.2%±1.2, Mean of 3 runs, range 38.2 to 40.73.9K$0.00968.68s
31low
58.8%
46.4%45.1%2.6K$0.001314.71s
32low
58.6%
±0.7, Mean of 3 runs, range 58.0 to 59.4
36.9%±0.8, Mean of 3 runs, range 36.1 to 37.735.8%±0.5, Mean of 3 runs, range 35.3 to 36.34.2K$0.01125.79s
33low
57.5%
46.6%44.5%2.0K$0.00234.09s
34low
57.1%
±1.7, Mean of 3 runs, range 55.9 to 59.4
48.2%±1.9, Mean of 3 runs, range 46.4 to 50.245.9%±1.6, Mean of 3 runs, range 44.5 to 47.72.1K$0.004118.59s
35low
57.0%
±1.3, Mean of 3 runs, range 56.1 to 58.7
45.9%±1.7, Mean of 3 runs, range 44.1 to 47.643.6%±1.0, Mean of 3 runs, range 42.6 to 44.53.8K$010.26s
36high
56.6%
±2.4, Mean of 3 runs, range 54.5 to 59.3
33.0%±2.3, Mean of 3 runs, range 30.9 to 35.433.5%±1.6, Mean of 3 runs, range 32.0 to 35.25.5K$0.01739.05s
37low
56.5%
39.9%39.2%3.3K$0.00528.59s
38low
56.4%
31.0%31.4%3.2K$0.05912.78s
39low
54.5%
43.2%41.2%3.8K$0.002415.64s
40
AnthropicClaude Opus 5
low
54.4%
30.2%30.5%3.1K$0.02711.20s
41low
54.3%
42.1%40.8%4.0K$092.14s
42low
53.9%
44.2%41.7%3.4K$0.002117.63s
43low
52.9%
±3.2, Mean of 3 runs, range 49.5 to 55.9
42.0%±2.2, Mean of 3 runs, range 39.8 to 44.140.3%±2.3, Mean of 3 runs, range 37.9 to 42.55.7K$024.78s
44low
52.1%
42.9%40.3%3.1K$0.001418.26s
45
MoonshotAIKimi K3
low
51.9%
41.5%40.4%4.3K$0.01823.23s
46low
50.5%
±3.5, Mean of 3 runs, range 46.1 to 53.0
39.7%±3.8, Mean of 3 runs, range 35.4 to 42.938.4%±3.2, Mean of 3 runs, range 34.4 to 40.85.5K$066.38s
47low
48.2%
±0.2, Mean of 3 runs, range 48.0 to 48.4
28.3%±0.7, Mean of 3 runs, range 27.8 to 29.229.9%±0.3, Mean of 3 runs, range 29.5 to 30.12.8K$045.25s
48low
46.6%
±0.6, Mean of 3 runs, range 45.8 to 47.0
36.0%±1.4, Mean of 3 runs, range 35.0 to 37.935.0%±0.9, Mean of 3 runs, range 34.4 to 36.16.2K$028.51s
49
DeepSeekDeepSeek V4 Flash Vision Exp
low
45.3%
28.8%29.9%1.1K$0.00117.05s
50
OpenAIGPT-5.5
high
44.2%
±0.6, Mean of 3 runs, range 43.5 to 44.8
24.0%±1.8, Mean of 3 runs, range 21.8 to 25.324.0%±0.8, Mean of 3 runs, range 23.0 to 24.66.0K$0.13053.30s
51low
44.2%
±0.7, Mean of 3 runs, range 43.5 to 44.8
29.9%±0.5, Mean of 3 runs, range 29.2 to 30.329.6%±0.5, Mean of 3 runs, range 29.3 to 30.33.8K$033.16s
52
OpenAIGPT-5.5
low
43.6%
±2.2, Mean of 3 runs, range 41.7 to 46.1
21.1%±1.4, Mean of 3 runs, range 19.8 to 22.622.5%±1.2, Mean of 3 runs, range 21.5 to 23.92.9K$0.03413.09s
53low
42.8%
32.9%31.1%2.2K$0.00019.54s
54low
41.0%
22.3%23.0%3.1K$0.002015.03s
55low
38.6%
29.2%27.9%2.0K$0.00315.86s
56low
38.6%
13.3%16.8%3.0K$0.0268.34s
57low
38.1%
±3.6, Mean of 3 runs, range 34.3 to 41.5
30.2%±1.6, Mean of 3 runs, range 28.3 to 31.528.9%±1.9, Mean of 3 runs, range 27.0 to 30.86.0K$021.74s
58low
36.1%
15.8%18.1%3.1K$0.0117.19s
59
MoonshotAIKimi K2.6
low
35.7%
26.7%26.8%3.5K$0.005964.70s
60low
35.0%
±1.5, Mean of 3 runs, range 33.2 to 36.1
28.6%±1.3, Mean of 3 runs, range 27.0 to 29.627.6%±1.2, Mean of 3 runs, range 26.1 to 28.44.1K$015.57s
61low
33.7%
25.1%24.7%1.4K$0.01010.06s
62low
33.1%
17.9%19.0%3.9K$0.00089.10s
63low
31.5%
±1.9, Mean of 3 runs, range 29.9 to 33.7
24.3%±1.6, Mean of 3 runs, range 22.9 to 26.223.6%±1.4, Mean of 3 runs, range 22.4 to 25.27.2K$018.83s
64low
23.8%
±1.1, Mean of 3 runs, range 22.7 to 24.9
16.6%±1.5, Mean of 3 runs, range 14.9 to 18.016.4%±1.1, Mean of 3 runs, range 15.3 to 17.41.1K$05.39s
65low
20.2%
4.6%8.0%2.3K$0.00686.47s
66low
19.0%
±1.2, Mean of 3 runs, range 18.0 to 20.4
12.7%±1.9, Mean of 3 runs, range 10.7 to 14.412.3%±1.1, Mean of 3 runs, range 11.3 to 13.5999$03.20s
67low
18.0%
4.6%7.5%3.0K$0.010018.73s
68high
16.6%
±0.8, Mean of 3 runs, range 15.8 to 17.4
3.0%±0.8, Mean of 3 runs, range 2.0 to 3.65.8%±0.1, Mean of 3 runs, range 5.7 to 5.98.2K$0.03042.45s
69low
15.8%
±0.4, Mean of 3 runs, range 15.3 to 16.1
3.3%±0.2, Mean of 3 runs, range 3.0 to 3.55.9%±0.2, Mean of 3 runs, range 5.7 to 6.12.5K$0.00447.32s
70low
0.0%
0.0%0.0%2.7K$020.86s

A ± after a score is half the range across that configuration's repeated runs; hover it for the run count and range. Scores without one are single runs. Every model is being re-run three times per task at each effort under the current protocol.

Score vs. cost

Object Detection score (mAP@50) against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

38 models on the current benchmark · Object Detection task only, low effort

Example Object Detection benchmark tasks

Real samples from the benchmark: the image each model sees, the question it is asked, and the ground-truth answer it is scored against.

Benchmark sample: Sports (Basketball) without annotations
Benchmark sample: Sports (Basketball) with predicted boxes and labels
AnnotatedOriginal

The models are asked

Detect all objects in this image. Output a JSON list where each entry contains the 2D bounding box in the key "box_2d" and the text label in the key "label". The "box_2d" value must be [y_min, x_min, y_max, x_max]: integers between 0 and 1000, normalized to the image height and width. Return only the JSON list, with no extra text. Only use these labels: ball in basket, basket rim, basketball, jersey number, player, player doing layup dunk, referee

ball in basketbasket rimbasketballjersey numberplayerplayer doing layup dunkreferee

Ground truth

19 objects

ball in basket × 1basket rim × 1basketball × 1jersey number × 5player × 8player doing layup dunk × 1referee × 2

Benchmark sample: Aerial / Drone (Solar Hot Spots) without annotations
Benchmark sample: Aerial / Drone (Solar Hot Spots) with predicted boxes and labels
AnnotatedOriginal

The models are asked

Detect all objects in this image. Output a JSON list where each entry contains the 2D bounding box in the key "box_2d" and the text label in the key "label". The "box_2d" value must be [y_min, x_min, y_max, x_max]: integers between 0 and 1000, normalized to the image height and width. Return only the JSON list, with no extra text. Only use these labels: solar panel hot spot

solar panel hot spot

Ground truth

15 objects

solar panel hot spot × 15

Benchmark sample: Aerial / Drone (Agriculture) without annotations
Benchmark sample: Aerial / Drone (Agriculture) with predicted boxes and labels
AnnotatedOriginal

The models are asked

Detect all objects in this image. Output a JSON list where each entry contains the 2D bounding box in the key "box_2d" and the text label in the key "label". The "box_2d" value must be [y_min, x_min, y_max, x_max]: integers between 0 and 1000, normalized to the image height and width. Return only the JSON list, with no extra text. Only use these labels: banana tree

banana tree

Ground truth

43 objects

banana tree × 43

Benchmark sample: Industrial (Bottle Caps) without annotations
Benchmark sample: Industrial (Bottle Caps) with predicted boxes and labels
AnnotatedOriginal

The models are asked

Detect all objects in this image. Output a JSON list where each entry contains the 2D bounding box in the key "box_2d" and the text label in the key "label". The "box_2d" value must be [y_min, x_min, y_max, x_max]: integers between 0 and 1000, normalized to the image height and width. Return only the JSON list, with no extra text. Only use these labels: soda bottle cap

soda bottle cap

Ground truth

24 objects

soda bottle cap × 24

Benchmark sample: Pharmaceutical (Pills) without annotations
Benchmark sample: Pharmaceutical (Pills) with predicted boxes and labels
AnnotatedOriginal

The models are asked

Detect all objects in this image. Output a JSON list where each entry contains the 2D bounding box in the key "box_2d" and the text label in the key "label". The "box_2d" value must be [y_min, x_min, y_max, x_max]: integers between 0 and 1000, normalized to the image height and width. Return only the JSON list, with no extra text. Only use these labels: pill

pill

Ground truth

50 objects

pill × 50

How Object Detection is scored

Quality is scored by how well the predicted boxes overlap the ground truth, using mean Average Precision. The headline metric is mAP@50 (a prediction counts when its box overlaps the truth by at least 50%), with stricter mAP@75 and mAP@50:95 also reported.

Every model runs the same sample set. Under the current protocol each task is run three times at each effort tier and the score is the mean over those runs, shown with its ± range in the table; models benchmarked before the protocol ran once per task and are being re-run under it. Token usage is measured from each provider’s API response, and cost per sample is that usage multiplied by the model’s published pricing. See the full methodology.

Frequently Asked Questions

Mean Average Precision measures how well predicted boxes match ground-truth boxes. The number is the overlap (IoU) threshold a prediction must clear to count: mAP@50 requires 50% overlap, mAP@75 requires 75%, and mAP@50:95 averages across thresholds from 50% to 95%. Higher thresholds demand more precise boxes, so scores drop as the threshold rises.

These are general-purpose vision language models prompted to output boxes, with no task-specific training. Dedicated detectors trained on a specific dataset still score far higher on it. The value here is zero-shot flexibility: the same model can detect arbitrary classes described in plain text.

Every task is run at two levels of reasoning effort. The low row is the model answering with minimal deliberation, the high row is the same model on the same questions allowed to think longer. Both rows are ranked together so you can see whether the extra thinking is worth its cost and latency, which the Est. cost and Speed columns show, and the effort filter above the table shows one tier at a time. Models benchmarked before the current protocol have a high-effort row for Reasoning only until they are re-run. Every other table on this site, including the overall Average, uses the low-effort pass so all models are compared on the same amount of compute.

Under the current protocol each task is run three times at each effort tier, and the score shown is the mean over those runs. The ± is half the range between the lowest and highest run, so it shows how much the score moved between runs; it is a range, not a confidence interval. Hover the ± for the run count and range. A score without one comes from a single run; models benchmarked before the protocol are being re-run under it.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.