Roboflow
Anthropic

Anthropic: Claude Opus 5.5

Not available to run yet

Claude Opus 5.5 is on the Vision Evals leaderboard. Running it in the Playground is not available yet. View evals

Claude Opus 5.5 Overview

Claude Opus 5.5 is a proprietary multimodal reasoning model from Anthropic and the first entry in the Claude 5.5 family. It accepts interleaved text and image input and returns text, with a one million token context window and up to 128,000 output tokens per response. Adaptive thinking is always enabled on this model and cannot be disabled; thinking depth is instead governed by an effort parameter with five levels, where medium is the default, a change from the high default used by Claude Opus 5 and earlier Opus models. Anthropic reports a knowledge cutoff of June 2026.

On the visual side, Anthropic characterizes Opus 5.5 as its strongest Opus release for vision and computer use, describing improved reading of dense documents, charts, screenshots, and diagrams for document extraction and visual analysis tasks. Published results include 89.0% on Chartography with tools and 81.8% on OSWorld 2.0 under partial credit scoring, alongside 48.7% under strict scoring reported in the system card. The accompanying system card states that Opus 5.5 scored higher than Opus 5 on every evaluation in its capability summary, with the largest gains concentrated in agentic coding, visual reasoning, computer use, and long-horizon knowledge work. The model ships with safety classifiers covering biology and cybersecurity that can route blocked requests to earlier Claude models.

Claude Opus 5.5 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRVision LanguageVisual Question AnsweringObject Detection

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Claude Opus 5.5 Vision Evals

Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.

Evals updated September 22, 2026Pricing updated September 22, 2026

Overall score#3 of 57
85.5%
Avg cost / sample#51 of 57
$0.014
Avg speed / sample#37 of 57
12.76s
Avg tokens / sample
2.2K

Strengths and weaknesses

Claude Opus 5.5 averages 85.5% across the six Vision Evals tasks, ranking #3 of 57 models overall.

It places in the top three for Counting, Reasoning, and Object Detection.

Its weakest relative showing is OCR, ranking #37 of 57 at 87.8%.

At $0.014 per sample it is the 51st cheapest of the 57 benchmarked models, and its average inference time of 12.8s per sample makes it the 37th fastest.

Performance profile

Field medianClaude Opus 5.5

Field medians: Object Detection 54.3%, Counting 61.7%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.8%.

Results by task

TaskScoreField (0 to 100)RankCost / sampleSpeed
Object Detection (low)
74.4%
±0.5, Mean of 3 runs, range 73.9 to 74.8
#3 of 57$0.02217.55s
Object Detection (high)
76.8%
±1.2, Mean of 3 runs, range 75.4 to 77.8
#3 of 23$0.03014.07s
Counting (low)
80.6%
±2.0, Mean of 3 runs, range 78.4 to 82.4
#2 of 57$0.008110.24s
Counting (high)
82.0%
±2.0, Mean of 3 runs, range 79.7 to 83.8
#2 of 23$0.009811.95s
Identification (low)
93.8%
±0.0, Mean of 3 runs, range 93.8 to 93.8
#8 of 57$0.00586.22s
Identification (high)
95.8%
±1.6, Mean of 3 runs, range 93.8 to 96.9
#6 of 23$0.00679.81s
OCR (low)
87.8%
±0.6, Mean of 3 runs, range 87.0 to 88.2
#37 of 57$0.0178.45s
OCR (high)
87.2%
±0.6, Mean of 3 runs, range 86.5 to 87.8
#22 of 23$0.02415.90s
Data Extraction (low)
93.5%
±0.5, Mean of 3 runs, range 92.8 to 93.8
#7 of 57$0.00668.80s
Data Extraction (high)
93.5%
±0.5, Mean of 3 runs, range 92.8 to 93.8
#5 of 23$0.007511.54s
Reasoning (low)
83.0%
±1.0, Mean of 3 runs, range 82.1 to 84.1
#2 of 57$0.009011.05s
Reasoning (high)
85.9%
±2.6, Mean of 3 runs, range 82.8 to 88.1
#2 of 43$0.0119.12s
  • Thinking longer helps: 2.4 points higher on object detection at high effort for 1.4x the cost and 0.8x the latency.
  • Thinking longer helps: 1.4 points higher on counting at high effort for 1.2x the cost and 1.2x the latency.
  • Thinking longer helps: 2.1 points higher on identification at high effort for 1.2x the cost and 1.6x the latency.
  • Thinking longer does not help: 0.6 points lower on ocr at high effort for 1.5x the cost and 1.9x the latency.
  • Thinking longer changes nothing: the same data extraction score at high effort for 1.1x the cost and 1.3x the latency.
  • Thinking longer helps: 2.9 points higher on reasoning at high effort for 1.2x the cost and 0.8x the latency.

Price vs. performance

Score vs. cost

Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.

56 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Claude Opus 5.5 highlighted

Claude Opus 5.5 scores are the mean of 3 runs per task at both low and high effort · Methodology

View all Vision Evals →

Claude Opus 5.5 Pricing

Claude Opus 5.5 costs $4.00 per 1M input tokens and $20.00 per 1M output tokens.

Input$4.00 / 1M tokens
Output$20.00 / 1M tokens
Cached input$0.200 / 1M tokens

Pricing updated Sep 22, 2026

Other Anthropic Opus models

Other versions in the same family as Claude Opus 5.5.

Claude Opus 5.5 License

Proprietary

Claude Opus 5.5 is proprietary: the weights are not distributed, and the Claude Opus 5.5 license is the vendor's commercial terms of service that you accept when you call the API.

Commercial use
Permitted under the vendor terms, typically metered per token or per request, with the vendor usage policy applying to your inputs and outputs.
Modification
Not available. Claude Opus 5.5 weights are closed, so you can configure prompts and use vendor-hosted fine-tuning where it is offered, but you cannot modify the model itself.
Redistribution
Not permitted. You cannot self-host or resell the model; you build on the hosted API instead.

Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.

Do I need a commercial license for Claude Opus 5.5?

Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Claude Opus 5.5.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Claude Opus 5.5 Vision

Yes. Claude Opus 5.5 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Counting at 80.6% (#2 of 57 at low effort).

Yes. its transcriptions match the ground truth 87.8% on average (#37 of 57 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 93.5%.

Yes. On Vision Evals, Claude Opus 5.5 scores 74.4% mAP@50 on object detection (#3 of 57 at low effort) and 80.6% judge-graded accuracy on object counting.

On our benchmark's task mix, Claude Opus 5.5 averages $0.01 per sample at $4.00 per 1M input and $20.00 per 1M output tokens (#51 of 57 on cost), with an average speed of 12.8s per sample across the benchmark. Actual cost depends on your images and prompts.

On the overall Vision Evals ranking, Claude Opus 5.5 sits #3 of 57 at 85.5%, just behind Gemini 3.5 Flash (86%) and just ahead of Gemini 3.7 Flash (85.2%). See the full side-by-side: Claude Opus 5.5 vs Gemini 3.5 Flash.