Roboflow

OCR Benchmark

The OCR task asks each model to read and transcribe every piece of text in an image, exactly as it appears. The test set spans printed documents, handwriting, code, receipts, forms, math, and hard-to-read strings such as serial numbers and license plates. Capitalization, punctuation, symbols, and layout must be preserved, and nothing may be summarized, corrected, or skipped.

59 models evaluated84 runs across effort levels

Evals updated September 22, 2026Pricing updated September 25, 2026

Score key:≥75%40–74%<40%
1low
94.0%
1.9K$0.03910.87s
2low
94.0%
±0.4, Mean of 3 runs, range 93.6 to 94.4
1.9K$0.0399.76s
3low
93.8%
1.9K$0.0208.27s
4high
93.6%
±0.2, Mean of 3 runs, range 93.5 to 93.9
1.9K$0.0399.66s
5low
93.6%
±0.6, Mean of 3 runs, range 92.9 to 94.1
2.8K$0.00797.37s
6
GrokGrok 4.7NEW
high
93.5%
±0.3, Mean of 3 runs, range 93.1 to 93.8
8.6K$0.03480.11s
7low
93.3%
±0.5, Mean of 3 runs, range 92.8 to 93.9
1.6K$0.005613.96s
8
AnthropicClaude Opus 5
low
93.2%
1.9K$0.0208.78s
9
MoonshotAIKimi K3
low
93.0%
1.7K$0.009413.18s
10high
92.9%
±0.8, Mean of 3 runs, range 91.9 to 93.6
4.2K$0.01414.79s
11high
92.8%
±0.3, Mean of 3 runs, range 92.6 to 93.2
4.3K$0.01410.40s
12low
92.6%
1.5K$0.00666.00s
13
GrokGrok 4.7NEW
low
92.6%
±0.7, Mean of 3 runs, range 92.1 to 93.4
4.6K$0.01430.02s
14low
92.6%
±0.5, Mean of 3 runs, range 92.1 to 93.1
2.4K$0.00636.51s
15high
92.3%
±0.5, Mean of 3 runs, range 91.9 to 92.9
3.0K$0.01228.16s
16low
92.2%
±1.2, Mean of 3 runs, range 91.1 to 93.4
1.8K$012.32s
17low
92.1%
1.9K$0.00107.66s
18low
92.1%
±0.3, Mean of 3 runs, range 91.9 to 92.5
2.2K$0.006818.03s
19low
91.9%
±0.2, Mean of 3 runs, range 91.6 to 92.1
1.6K$0.0317.02s
19
OpenAIGPT-6 SolNEW
high
91.9%
±0.3, Mean of 3 runs, range 91.6 to 92.2
2.9K$0.01919.70s
21low
91.8%
±0.3, Mean of 3 runs, range 91.5 to 92.1
2.6K$0.009112.87s
22low
91.7%
1.9K$0.00787.23s
22
OpenAIGPT-6 SolNEW
low
91.7%
±0.4, Mean of 3 runs, range 91.3 to 92.1
1.7K$0.00707.25s
24
OpenAIGPT-5.5
high
91.7%
±0.6, Mean of 3 runs, range 91.1 to 92.3
3.4K$0.07226.26s
25high
91.6%
±0.2, Mean of 3 runs, range 91.4 to 91.7
5.0K$0.02348.72s
26high
91.5%
±0.2, Mean of 3 runs, range 91.3 to 91.7
2.8K$0.08925.89s
27high
91.5%
±1.4, Mean of 3 runs, range 90.1 to 92.9
2.0K$014.10s
28high
91.5%
±0.3, Mean of 3 runs, range 91.2 to 91.7
4.6K$0.004230.81s
29high
91.3%
±0.5, Mean of 3 runs, range 90.8 to 91.9
1.8K$0.00059.12s
30high
91.3%
±0.5, Mean of 3 runs, range 90.7 to 91.7
5.2K$0.02789.51s
31low
91.3%
±0.5, Mean of 3 runs, range 90.7 to 91.6
2.9K$0.008332.47s
32
OpenAIGPT-5.5
low
91.2%
±0.3, Mean of 3 runs, range 90.9 to 91.6
1.8K$0.0249.53s
33low
90.8%
±0.6, Mean of 3 runs, range 90.2 to 91.5
3.5K$067.50s
34low
90.7%
±1.7, Mean of 3 runs, range 88.5 to 91.9
1.4K$0.00088.11s
35low
90.7%
±1.8, Mean of 3 runs, range 88.4 to 92.0
2.0K$0.00127.52s
35low
90.7%
±0.1, Mean of 3 runs, range 90.6 to 90.7
2.1K$0.01111.86s
37low
90.6%
1.7K$0.00017.41s
38high
90.2%
±0.2, Mean of 3 runs, range 90.0 to 90.4
3.5K$0.02530.98s
39low
89.5%
±1.3, Mean of 3 runs, range 88.1 to 90.6
2.0K$0.00426.66s
40high
89.5%
±0.0, Mean of 3 runs, range 89.5 to 89.6
5.5K$0.01724.69s
41high
89.4%
±0.6, Mean of 3 runs, range 88.8 to 90.1
3.0K$0.02323.00s
42low
89.4%
±0.8, Mean of 3 runs, range 88.8 to 90.3
2.1K$0.01211.59s
43low
89.3%
±1.6, Mean of 3 runs, range 88.0 to 91.1
2.8K$0.0169.19s
44low
89.3%
1.7K$0.003012.50s
45high
89.0%
±0.8, Mean of 3 runs, range 88.3 to 89.9
3.4K$0.00939.60s
46high
89.0%
±1.8, Mean of 3 runs, range 87.7 to 91.2
7.6K$0.03034.24s
47high
88.9%
±0.2, Mean of 3 runs, range 88.7 to 89.1
4.8K$0.03516.38s
48high
88.8%
±0.7, Mean of 3 runs, range 88.0 to 89.4
17.9K$0.06457.30s
49low
88.8%
748$0.00475.17s
50low
88.7%
±1.3, Mean of 3 runs, range 87.6 to 90.2
6.8K$064.46s
51low
88.5%
±1.9, Mean of 3 runs, range 86.7 to 90.6
3.1K$036.22s
52
OpenAIGPT-6 LunaNEW
high
88.5%
±0.6, Mean of 3 runs, range 87.9 to 89.2
3.8K$0.001422.29s
53
MoonshotAIKimi K2.6
low
88.5%
1.7K$0.002116.21s
54low
88.2%
±1.6, Mean of 3 runs, range 86.9 to 90.0
1.6K$0.002711.56s
55low
88.2%
±0.3, Mean of 3 runs, range 87.9 to 88.4
1.7K$0.00286.33s
56low
88.1%
1.5K$0.001010.50s
57low
88.0%
±1.0, Mean of 3 runs, range 86.8 to 88.9
1.4K$0.00036.70s
58
OpenAIGPT-6 LunaNEW
low
87.9%
±0.6, Mean of 3 runs, range 87.2 to 88.3
2.0K$0.00058.59s
59
AnthropicClaude Opus 5.5NEW
low
87.8%
±0.6, Mean of 3 runs, range 87.0 to 88.2
2.0K$0.0178.45s
60low
87.7%
±0.0, Mean of 3 runs, range 87.6 to 87.7
2.9K$019.73s
61low
87.6%
1.8K$0.00244.21s
62high
87.5%
±2.7, Mean of 3 runs, range 85.3 to 90.6
6.0K$0.0048101.54s
63low
87.4%
1.5K$0.00111.57s
64low
87.3%
±0.8, Mean of 3 runs, range 86.5 to 88.2
1.5K$0.00245.58s
65
AnthropicClaude Opus 5.5NEW
high
87.2%
±0.6, Mean of 3 runs, range 86.5 to 87.8
2.3K$0.02415.90s
66low
87.0%
±0.5, Mean of 3 runs, range 86.6 to 87.7
1.5K$0.000315.58s
67high
87.0%
±2.1, Mean of 3 runs, range 84.3 to 88.5
6.3K$0.001673.66s
68high
86.9%
±4.1, Mean of 3 runs, range 82.2 to 90.4
4.4K$0.01557.96s
69low
86.7%
±1.6, Mean of 3 runs, range 85.6 to 88.8
2.9K$015.73s
70
DeepSeekDeepSeek V4 Flash Vision Exp
low
86.7%
771$0.00074.42s
71low
86.5%
1.5K$0.000910.39s
72low
84.7%
±3.3, Mean of 3 runs, range 80.8 to 87.3
5.4K$0106.58s
73low
84.5%
1.5K$0.00098.24s
74low
84.2%
±0.9, Mean of 3 runs, range 83.0 to 84.9
7.0K$052.61s
75low
84.1%
1.5K$0.000111.39s
76low
83.1%
±1.9, Mean of 3 runs, range 81.8 to 85.5
1.2K$08.12s
77low
83.0%
±0.4, Mean of 3 runs, range 82.7 to 83.5
6.1K$041.09s
78low
82.8%
1.5K$0.001616.63s
79low
81.4%
1.5K$0.001810.57s
80low
81.1%
1.2K$030.31s
81low
80.4%
±0.7, Mean of 3 runs, range 79.9 to 81.2
1.4K$07.32s
82low
80.2%
±4.9, Mean of 3 runs, range 75.4 to 85.1
7.3K$043.05s
83low
66.7%
±5.0, Mean of 3 runs, range 62.9 to 72.9
9.2K$043.70s
84low
59.5%
±2.2, Mean of 3 runs, range 57.6 to 62.0
1.5K$04.17s

A ± after a score is half the range across that configuration's repeated runs; hover it for the run count and range. Scores without one are single runs. Every model is being re-run three times per task at each effort under the current protocol.

Score vs. cost

OCR score (Mean Similarity) against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

44 models on the current benchmark · OCR task only, low effort

Example OCR benchmark tasks

Real samples from the benchmark: the image each model sees, the question it is asked, and the ground-truth answer it is scored against.

Benchmark sample: Handwriting (Recipe)

The models are asked

Transcribe all handwritten and printed text exactly as written, reading top to bottom and left to right, and omit nothing; mark only genuinely unreadable words as [illegible]. Output plain text only, with no Markdown or HTML, and use a single space wherever text is indented or aligned.

Ground truth

100 Good Cookies 1 cup white sugar 1 teaspoon baking soda 1 cup brown sugar 1 cup quick oatmeal 1 cup margarine 1 cup coconut 1 cup cooking oil 1 cup Rice Krispies 1 egg 1 cup chopped walnuts 1 teaspoon vanilla 3½ cups flour 1 teaspoon cream of tartar 1-12oz. package mini chocolate chips Mix in order given. Drop by teaspoonful onto cookie sheet. Bake at 350° for 8-10 minutes. (over)

Benchmark sample: Printed Table (Cost Of Living)

The models are asked

Transcribe all text exactly as written, reading top to bottom and left to right, and omit nothing. Output plain text only, with no Markdown or HTML, and use a single space wherever text is indented or aligned.

Ground truth

COST OF LIVING The following figures are a general guide to essential monthly expenses for an expatriate executive living in Azerbaijan. Married Single couple with person two children (US$) (US$) Rent 800 3,000 Food 300 1,200 Transport 150 300 Utilities 100 150 Education n. a. 1,000 Miscellaneous 500 1,000 Total 1,850 6,650

Benchmark sample: Invoice With Table (HTML Structure)

The models are asked

Transcribe all text in this document exactly and completely, in reading order, preserving casing, punctuation, currency symbols, and codes. Use an HTML table only for genuinely tabular content, mirroring its rows, columns, and merged cells; write the HTML on single lines with no indentation or blank lines between tags, and collapse horizontal alignment to a single space.

Ground truth

TOM GREEN HANDYMAN 5 Any Street, Any City, That Area Code Telephone: 0800 XXX XXX Date : 6/5/2016 Invoice No : 0003521 Tax Registered No 123456 Mr and Mrs Fielding This Address This City This Area Code TAX INVOICE <table><tr><th>Quantity</th><th>Description</th><th>Unit Price</th><th>Cost</th></tr><tr><td></td><td>Upgrade to Bathroom</td><td></td><td></td></tr><tr><td>23.75</td><td>Labour</td><td>40.00</td><td>950.00</td></tr><tr><td>50</td><td>Nails and screws</td><td>0.80</td><td>40.00</td></tr><tr><td>1</td><td>Paint and Plywood</td><td>1000.00</td><td>1000.00</td></tr><tr><td>40</td><td>Imported wall tiles</td><td>14.00</td><td>560.00</td></tr><tr><td>1</td><td>Freight</td><td>150.00</td><td>150.00</td></tr><tr><td>1</td><td>Sub-contractor : Tile-It</td><td></td><td>228.00</td></tr><tr><td></td><td></td><td>Subtotal</td><td>2928.00</td></tr><tr><td></td><td></td><td>Tax</td><td>439.20</td></tr><tr><td></td><td></td><td>Total Due</td><td>$3,367.20</td></tr></table> Payment due by the 10th of the month following the date of invoice. Please make payment into Bank Account No. 12 3456 789112 012 Interest of 10% per year will be charged on late payments. Cut here Remittance Mr and Mrs Fielding TOM GREEN HANDYMAN 5 Any Street Any City That Area Code Amount Due $3,367.20 Amount Paid ______

Benchmark sample: Math Exam (LaTeX)

The models are asked

Transcribe every piece of text in this image exactly as written, in reading order, preserving question and option numbering, and omit nothing. Write mathematical expressions in LaTeX and format the result as Markdown that mirrors the layout.

Ground truth

Q.5 Let $L_1$ be the line of intersection of the planes given by the equations $$2x + 3y + z = 4 \quad \text{and} \quad x + 2y + z = 5.$$ Let $L_2$ be the line passing through the point $P(2, -1, 3)$ and parallel to $L_1$. Let $M$ denote the plane given by the equation $$2x + y - 2z = 6.$$ Suppose that the line $L_2$ meets the plane $M$ at the point $Q$. Let $R$ be the foot of the perpendicular drawn from $P$ to the plane $M$ . Then which of the following statements is (are) TRUE? | | | |---|---| | (A) | The length of the line segment $PQ$ is $9\sqrt{3}$ | | (B) | The length of the line segment $QR$ is $15$ | | (C) | The area of $\triangle PQR$ is $\dfrac{3}{2}\sqrt{234}$ | | (D) | The acute angle between the line segments $PQ$ and $PR$ is $\cos^{-1}\left(\dfrac{1}{2\sqrt{3}}\right)$ | Answer: A, C 3/9

Benchmark sample: Social Media Post

The models are asked

Transcribe all text exactly as written, reading top to bottom, keeping each author's handle with their message, and omit nothing. Output plain text only, with no Markdown or HTML, and use a single space wherever text is indented or aligned. Transcribe only the first tweet by kache (@yacineMTB).

Ground truth

I'm losing so much money every single second I dont spend in the states. Every day I wake up, I'm hyperventilating about the fact that I am irresponsibly putting my young family's future at great risk because I can't but help listen to my gut. I am filled with doubt every day

How OCR is scored

Each transcription is scored by its similarity to the ground-truth text, so partial credit is possible. The leaderboard score is the mean similarity across all samples, expressed as a percentage.

Every model runs the same sample set. Under the current protocol each task is run three times at each effort tier and the score is the mean over those runs, shown with its ± range in the table; models benchmarked before the protocol ran once per task and are being re-run under it. Token usage is measured from each provider’s API response, and cost per sample is that usage multiplied by the model’s published pricing. See the full methodology.

Frequently Asked Questions

Each model transcribes the same set of images. Every transcription is compared to the ground-truth text and scored by similarity, so a near-perfect answer earns most of the credit rather than failing outright. The leaderboard shows the mean similarity across all samples.

Printed documents, handwriting, code, receipts, forms, tables, math notation, and short high-stakes strings like serial numbers and license plates. Prompts require exact transcription with layout, casing, and punctuation preserved.

Every task is run at two levels of reasoning effort. The low row is the model answering with minimal deliberation, the high row is the same model on the same questions allowed to think longer. Both rows are ranked together so you can see whether the extra thinking is worth its cost and latency, which the Est. cost and Speed columns show, and the effort filter above the table shows one tier at a time. Models benchmarked before the current protocol have a high-effort row for Reasoning only until they are re-run. Every other table on this site, including the overall Average, uses the low-effort pass so all models are compared on the same amount of compute.

Under the current protocol each task is run three times at each effort tier, and the score shown is the mean over those runs. The ± is half the range between the lowest and highest run, so it shows how much the score moved between runs; it is a range, not a confidence interval. Hover the ± for the run count and range. A score without one comes from a single run; models benchmarked before the protocol are being re-run under it.

Most models in this leaderboard link to their Playground page. Click the model name to open it, then upload your own image and run it. A few models are benchmarked for comparison only and do not have a Playground page yet.