Roboflow

Data Extraction Benchmark

The Data Extraction task asks each model to find and return one specific piece of information from an image, rather than all of its text. The target is usually a single field on a structured or semi-structured source: a price, date, timestamp, phone number, license plate, serial number, or meter reading. The model must locate the right field among many competing ones and return just that value in the requested format.

34 models evaluated

Evals updated August 27, 2026Pricing updated August 31, 2026

Score key:≥75%40–74%<40%
1
96.9%
1.2K$0.00081.98s
2
94.8%
1.3K$0.00152.48s
2
94.8%
1.4K$0.00372.67s
2
94.8%
1.5K$0.00635.64s
2
94.8%
1.3K$0.00148.73s
6
92.8%
1.5K$0.0154.98s
7
90.7%
1.2K$0.00041.33s
8
89.7%
1.5K$0.00302.81s
9
AnthropicClaude Opus 5
88.7%
1.5K$0.00804.17s
9
88.7%
1.8K$0.00334.18s
11
87.6%
1.5K$0.00762.48s
11
OpenAIGPT-5.5
87.6%
1.5K$0.0104.74s
11
87.6%
1.2K$0.00295.59s
14
86.6%
1.6K$0.00062.17s
14
86.6%
1.1K$0.00022.85s
14
86.6%
1.1K$0.00024.58s
14
86.6%
1.8K$0.00314.63s
18
85.6%
1.1K$0.00072.02s
19
84.5%
400$0.00133.02s
19
84.5%
1.6K$0.00424.51s
19
MoonshotAIKimi K3
84.5%
1.5K$0.00465.65s
22
83.5%
1.1K$0.00042.42s
22
83.5%
1.5K$0.00014.78s
22
83.5%
1.8K$0.00445.18s
25
82.5%
1.4K$0.00142.84s
25
82.5%
1.4K$0.00333.42s
27
81.4%
1.5K$0.00042.94s
27
81.4%
1.4K$0.00184.35s
29
79.4%
1.1K$0.00051.24s
29
79.4%
1.4K$0.00373.06s
31
78.3%
1.1K$0.00022.20s
31
78.3%
1.1K<$0.00013.49s
31
MoonshotAIKimi K2.6
78.3%
1.6K$0.00157.14s
34
DeepSeekDeepSeek V4 Flash Vision Exp
65.0%
386$0.00021.26s

Score vs. cost

Data Extraction score (Accuracy) against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

34 models on the current benchmark · Data Extraction task only

Example Data Extraction benchmark tasks

Real samples from the benchmark: the image each model sees, the question it is asked, and the ground-truth answer it is scored against.

Benchmark sample: License Plate

The models are asked

What is the license plate number on the rear of the black SUV? Format the answer as a continuous sequence of uppercase letters and digits (e.g., ABC1234).

Ground truth

9CGN302

Benchmark sample: Container Serial Number

The models are asked

What is the complete two-line identification text printed in white on the upper right side of the container doors? Format the answer with a single space separating the top sequence from the bottom sequence.

Ground truth

TCKU 612608 6 45G1

Benchmark sample: Tire Sidewall Spec

The models are asked

What is the large alphanumeric tire size sequence printed on the right side of the sidewall? Output the exact text including spaces and slashes.

Ground truth

205/60 R 16

How Data Extraction is scored

Each answer is graded against the ground-truth value. The leaderboard score is the percentage of samples answered correctly.

Every model runs the same sample set in a single evaluation pass. Token usage is measured from each provider’s API response, and cost per sample is that usage multiplied by the model’s published pricing. See the full methodology.

Frequently Asked Questions

OCR asks for a complete transcription of everything in the image. Data Extraction asks for exactly one value, so the model must find the right field among many competing ones and return it in the requested format. A model can be strong at full-page OCR and still fail at targeted extraction.

Each model answers the same extraction prompts. Answers are graded against the ground truth, and the score is the percentage answered correctly.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.