Compare the best 69 OCR models and try 64 of them on your own image, free in the Roboflow Playground. 29 are open-weight, so you can self-host them for free under their licenses.
69 models · 29 open-weight · 64 free to try · prices synced Jul 28, 2026
The best OCR model on our OCR benchmark right now is Claude Fable 5 by Anthropic, scoring 94.0% across 23 models tested, followed by Claude Opus 4.8 at 93.8%. Updated Jul 28, 2026. See the full ranking.
Benchmark Leader
Claude Fable 5
Highest score on our OCR benchmark, 23 models tested
Open-Weight Leader
Kimi K3
Best score among the open-weight models, #4 of 23 overall on our OCR benchmark
Fastest on Benchmark
Gemini 3.5 Flash-Lite
Quickest in our OCR benchmark runs: 1.6s per request on average
Lowest Measured Cost
Qwen 3.7 Flash
Cheapest in our OCR benchmark runs: $0.0001 per request as measured, not list price
Ranked on our OCR benchmark, 23 models tested, updated Jul 28, 2026
29 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 24 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle NEW | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
GLM-OCRZ.ai | 0.9B | MIT | Mar 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
SuryaMindee | — | GPL v3 | Jan 2024 | |
TrOCRMicrosoft | 61.4M-600M | MIT | Sep 2021 | |
docTRMindee | — | Apache 2.0 | Feb 2021 |
40 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 | |
Google Vision OCRGoogle | — | — | — | Feb 2016 |
The right OCR model depends on what your documents look like and what you need back: raw text, structured fields, or answers. Clean printed pages, handwritten forms, and dense multi-column tables are different problems, and the models on this page split into two families that handle them differently.
Purpose-built OCR systems, from classic engines to OCR-tuned models like GLM-OCR, are optimized for one job: reading text accurately and cheaply at volume. They tend to win on cost per page, throughput, and consistency for well-understood document types, and specialist models increasingly handle tables, formulas, and layout structure that used to require separate systems.
Choose this family when you process documents at scale (invoices, receipts, archives), the format is broadly predictable, and cost per page matters more than open-ended flexibility.
Frontier VLMs read text as one part of understanding the whole page: they can transcribe, then also summarize, answer questions, or return the content already structured as JSON or markdown from a single prompt. That flexibility comes at a higher price per page and more latency, and on straightforward transcription a dedicated engine is usually cheaper for the same accuracy.
Choose a VLM when the extraction is complicated (mixed layouts, judgment calls, follow-up questions), when you want structured output without building a parsing pipeline, or when volumes are small enough that flexibility beats unit cost.
Teams processing documents at scale often route pages: a fast, cheap OCR model handles the bulk, and hard pages (poor scans, handwriting, unusual layouts) escalate to a stronger VLM. Confidence scores or simple heuristics decide the split, keeping average cost low without giving up accuracy on the tail.
The bottom line: High volume and predictable documents point to a dedicated OCR model; complex extraction and structured output point to a VLM. At scale, route easy pages cheap and escalate the hard ones.
OCR (optical character recognition) is the task of reading text out of images: scans, photos, screenshots, and documents. Classic pipelines split it into text detection (finding where text is) and text recognition (decoding the characters); modern vision language models read end to end and can return text already structured as markdown, JSON, or key-value pairs. Current models handle handwriting, rotated and curved text, tables, and multi-column layouts, which used to require separate systems. Quality is typically measured by character and word error rate (CER and WER) on benchmark document sets. OCR is the backbone of invoice and receipt processing, document digitization, identity verification, and search over scanned archives. This page lists 69 OCR models, including 29 open-weight options you can self-host; 64 of them run live in the Playground so you can test them on your own images.
On our OCR benchmark (23 models tested, updated Jul 28, 2026), Claude Fable 5 by Anthropic currently scores highest at 94.0%. The full ranking is on our evals page. For fixed categories in production, a model fine-tuned on your own data still often wins.
Yes. 29 of the 69 OCR models here are open-weight (for example Kimi K3, Gemma 4 12B, and Qwen3.6 27B), free to self-host under their licenses (Modified MIT, Apache 2.0, and MIT).
Yes. You can run 64 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 69 OCR models in the Roboflow Playground catalog: 29 open-weight models you can self-host and 40 proprietary models accessed through provider APIs; 64 of them run live in the Roboflow Playground on your own images. On our OCR benchmark (23 models tested), Claude Fable 5 by Anthropic currently scores highest. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.