Not available to run yet
Claude Sonnet 5.5 is on the Vision Evals leaderboard. Running it in the Playground is not available yet. View evals
Claude Sonnet 5.5 is a proprietary multimodal language model from Anthropic and the second release in the Claude 5.5 family, following Claude Opus 5.5. It accepts interleaved text and image input and returns text, operating with a 1M token context window and a maximum output of 128K tokens per request. The model uses adaptive thinking by default, allocating variable reasoning effort per request rather than exposing a manual extended thinking toggle, and its training data cutoff is June 2026. Anthropic positions it as a faster, lower cost complement to Opus 5.5 for well scoped everyday tasks, bug fixing, and producing documents, slides, and spreadsheets.
On visual and agentic evaluations reported at launch, Sonnet 5.5 scores 61.6% on Chartography, a chart recognition test, compared with 15.6% for Claude Sonnet 5, and 80.1% on OSWorld 2.1, a computer use benchmark measuring screenshot driven control of a desktop environment, compared with 57.0% for Sonnet 5. It reports 70.6% on Terminal-Bench 4.0 for agentic coding. Anthropic describes it as the first Sonnet model able to complete Pokemon Red from screenshots alone, and it generates output more than 30% faster than Sonnet 5 while using fewer tokens for equivalent work.
—
Usage
Past 30 DaysNot available
Not in Playground
Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated September 28, 2026Pricing updated September 28, 2026
Claude Sonnet 5.5 averages 83.8% across the six Vision Evals tasks, ranking #7 of 60 models overall.
Its weakest relative showing is OCR, ranking #26 of 60 at 90.6%.
At $0.0065 per sample it is the 41st cheapest of the 60 benchmarked models, and its average inference time of 10.8s per sample makes it the 36th fastest.
Field medians: Object Detection 54.1%, Counting 60.6%, Identification 84.4%, OCR 89.1%, Data Extraction 84.5%, Reasoning 55.9%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection (low) | 74.3% ±0.9, Mean of 3 runs, range 73.5 to 75.3 | #5 of 60 | $0.0098 | 13.60s | |
| Object Detection (high) | 76.8% ±0.4, Mean of 3 runs, range 76.5 to 77.3 | #4 of 26 | $0.014 | 9.81s | |
| Counting (low) | 79.3% ±0.7, Mean of 3 runs, range 78.4 to 79.7 | #6 of 60 | $0.0042 | 10.67s | |
| Counting (high) | 82.9% ±1.4, Mean of 3 runs, range 81.1 to 83.8 | #1 of 26 | $0.0053 | 10.91s | |
| Identification (low) | 91.7% ±3.1, Mean of 3 runs, range 87.5 to 93.8 | #13 of 60 | $0.0029 | 6.51s | |
| Identification (high) | 90.6% ±0.0, Mean of 3 runs, range 90.6 to 90.6 | #10 of 26 | $0.0033 | 5.96s | |
| OCR (low) | 90.6% ±0.9, Mean of 3 runs, range 90.0 to 91.7 | #26 of 60 | $0.0079 | 5.64s | |
| OCR (high) | 90.9% ±1.5, Mean of 3 runs, range 89.2 to 92.3 | #15 of 26 | $0.011 | 8.77s | |
| Data Extraction (low) | 90.7% ±1.5, Mean of 3 runs, range 89.7 to 92.8 | #11 of 60 | $0.0033 | 7.81s | |
| Data Extraction (high) | 93.1% ±0.5, Mean of 3 runs, range 92.8 to 93.8 | #7 of 26 | $0.0036 | 8.36s | |
| Reasoning (low) | 76.4% ±0.7, Mean of 3 runs, range 75.5 to 76.8 | #7 of 60 | $0.0049 | 10.22s | |
| Reasoning (high) | 83.9% ±1.7, Mean of 3 runs, range 82.1 to 85.4 | #4 of 46 | $0.0061 | 7.47s |
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
59 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Claude Sonnet 5.5 highlighted
Claude Sonnet 5.5 scores are the mean of 3 runs per task at both low and high effort · Methodology
View all Vision Evals →Other versions in the same family as Claude Sonnet 5.5.
Claude Sonnet 5.5 is proprietary: the weights are not distributed, and the Claude Sonnet 5.5 license is the vendor's commercial terms of service that you accept when you call the API.
Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.
Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Claude Sonnet 5.5.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. Claude Sonnet 5.5 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Object Detection at 74.3% (#5 of 60 at low effort).
Yes. its transcriptions match the ground truth 90.6% on average (#26 of 60 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 90.7%.
Yes. On Vision Evals, Claude Sonnet 5.5 scores 74.3% mAP@50 on object detection (#5 of 60 at low effort) and 79.3% judge-graded accuracy on object counting.
On our benchmark's task mix, Claude Sonnet 5.5 averages $0.0065 per sample (#41 of 60 on cost), with an average speed of 10.8s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Claude Sonnet 5.5 sits #7 of 60 at 83.8%, just behind Qwen3.8 Max (83.9%) and just ahead of Gemini 3.1 Pro (83.3%). See the full side-by-side: Claude Sonnet 5.5 vs Qwen3.8 Max.