Roboflow
Anthropic

Anthropic: Claude Sonnet 5.5

Not available to run yet

Claude Sonnet 5.5 is on the Vision Evals leaderboard. Running it in the Playground is not available yet. View evals

Claude Sonnet 5.5 Overview

Claude Sonnet 5.5 is a proprietary multimodal language model from Anthropic and the second release in the Claude 5.5 family, following Claude Opus 5.5. It accepts interleaved text and image input and returns text, operating with a 1M token context window and a maximum output of 128K tokens per request. The model uses adaptive thinking by default, allocating variable reasoning effort per request rather than exposing a manual extended thinking toggle, and its training data cutoff is June 2026. Anthropic positions it as a faster, lower cost complement to Opus 5.5 for well scoped everyday tasks, bug fixing, and producing documents, slides, and spreadsheets.

On visual and agentic evaluations reported at launch, Sonnet 5.5 scores 61.6% on Chartography, a chart recognition test, compared with 15.6% for Claude Sonnet 5, and 80.1% on OSWorld 2.1, a computer use benchmark measuring screenshot driven control of a desktop environment, compared with 57.0% for Sonnet 5. It reports 70.6% on Terminal-Bench 4.0 for agentic coding. Anthropic describes it as the first Sonnet model able to complete Pokemon Red from screenshots alone, and it generates output more than 30% faster than Sonnet 5 while using fewer tokens for equivalent work.

Claude Sonnet 5.5 Details & Performance

Details

Resources

—

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question AnsweringObject Detection

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Claude Sonnet 5.5 Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated September 28, 2026Pricing updated September 28, 2026

Overall score#7 of 60
83.8%
Avg cost / sample#41 of 60
$0.0065
Avg speed / sample#36 of 60
10.78s
Avg tokens / sample
2.1K

Strengths and weaknesses

Claude Sonnet 5.5 averages 83.8% across the six Vision Evals tasks, ranking #7 of 60 models overall.

Its weakest relative showing is OCR, ranking #26 of 60 at 90.6%.

At $0.0065 per sample it is the 41st cheapest of the 60 benchmarked models, and its average inference time of 10.8s per sample makes it the 36th fastest.

Performance profile

Field medianClaude Sonnet 5.5

Field medians: Object Detection 54.1%, Counting 60.6%, Identification 84.4%, OCR 89.1%, Data Extraction 84.5%, Reasoning 55.9%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection (low)
74.3%
±0.9, Mean of 3 runs, range 73.5 to 75.3
#5 of 60$0.009813.60s
Object Detection (high)
76.8%
±0.4, Mean of 3 runs, range 76.5 to 77.3
#4 of 26$0.0149.81s
Counting (low)
79.3%
±0.7, Mean of 3 runs, range 78.4 to 79.7
#6 of 60$0.004210.67s
Counting (high)
82.9%
±1.4, Mean of 3 runs, range 81.1 to 83.8
#1 of 26$0.005310.91s
Identification (low)
91.7%
±3.1, Mean of 3 runs, range 87.5 to 93.8
#13 of 60$0.00296.51s
Identification (high)
90.6%
±0.0, Mean of 3 runs, range 90.6 to 90.6
#10 of 26$0.00335.96s
OCR (low)
90.6%
±0.9, Mean of 3 runs, range 90.0 to 91.7
#26 of 60$0.00795.64s
OCR (high)
90.9%
±1.5, Mean of 3 runs, range 89.2 to 92.3
#15 of 26$0.0118.77s
Data Extraction (low)
90.7%
±1.5, Mean of 3 runs, range 89.7 to 92.8
#11 of 60$0.00337.81s
Data Extraction (high)
93.1%
±0.5, Mean of 3 runs, range 92.8 to 93.8
#7 of 26$0.00368.36s
Reasoning (low)
76.4%
±0.7, Mean of 3 runs, range 75.5 to 76.8
#7 of 60$0.004910.22s
Reasoning (high)
83.9%
±1.7, Mean of 3 runs, range 82.1 to 85.4
#4 of 46$0.00617.47s
  • Thinking longer helps: 2.5 points higher on object detection at high effort for 1.4x the cost and 0.7x the latency.
  • Thinking longer helps: 3.6 points higher on counting at high effort for 1.3x the cost and 1x the latency.
  • Thinking longer does not help: 1 points lower on identification at high effort for 1.1x the cost and 0.9x the latency.
  • Thinking longer helps: 0.3 points higher on ocr at high effort for 1.4x the cost and 1.6x the latency.
  • Thinking longer helps: 2.4 points higher on data extraction at high effort for 1.1x the cost and 1.1x the latency.
  • Thinking longer helps: 7.5 points higher on reasoning at high effort for 1.3x the cost and 0.7x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

59 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Claude Sonnet 5.5 highlighted

Claude Sonnet 5.5 scores are the mean of 3 runs per task at both low and high effort · Methodology

View all Vision Evals →

Other Anthropic Sonnet models

Other versions in the same family as Claude Sonnet 5.5.

Claude Sonnet 5.5 License

Proprietary

Claude Sonnet 5.5 is proprietary: the weights are not distributed, and the Claude Sonnet 5.5 license is the vendor's commercial terms of service that you accept when you call the API.

Commercial use
Permitted under the vendor terms, typically metered per token or per request, with the vendor usage policy applying to your inputs and outputs.
Modification
Not available. Claude Sonnet 5.5 weights are closed, so you can configure prompts and use vendor-hosted fine-tuning where it is offered, but you cannot modify the model itself.
Redistribution
Not permitted. You cannot self-host or resell the model; you build on the hosted API instead.

Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.

Do I need a commercial license for Claude Sonnet 5.5?

Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Claude Sonnet 5.5.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Claude Sonnet 5.5 Vision

Yes. Claude Sonnet 5.5 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Object Detection at 74.3% (#5 of 60 at low effort).

Yes. its transcriptions match the ground truth 90.6% on average (#26 of 60 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 90.7%.

Yes. On Vision Evals, Claude Sonnet 5.5 scores 74.3% mAP@50 on object detection (#5 of 60 at low effort) and 79.3% judge-graded accuracy on object counting.

On our benchmark's task mix, Claude Sonnet 5.5 averages $0.0065 per sample (#41 of 60 on cost), with an average speed of 10.8s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Claude Sonnet 5.5 sits #7 of 60 at 83.8%, just behind Qwen3.8 Max (83.9%) and just ahead of Gemini 3.1 Pro (83.3%). See the full side-by-side: Claude Sonnet 5.5 vs Qwen3.8 Max.