Not available to run yet
Claude Opus 5.5 is on the Vision Evals leaderboard. Running it in the Playground is not available yet. View evals
Claude Opus 5.5 is a proprietary multimodal reasoning model from Anthropic and the first entry in the Claude 5.5 family. It accepts interleaved text and image input and returns text, with a one million token context window and up to 128,000 output tokens per response. Adaptive thinking is always enabled on this model and cannot be disabled; thinking depth is instead governed by an effort parameter with five levels, where medium is the default, a change from the high default used by Claude Opus 5 and earlier Opus models. Anthropic reports a knowledge cutoff of June 2026.
On the visual side, Anthropic characterizes Opus 5.5 as its strongest Opus release for vision and computer use, describing improved reading of dense documents, charts, screenshots, and diagrams for document extraction and visual analysis tasks. Published results include 89.0% on Chartography with tools and 81.8% on OSWorld 2.0 under partial credit scoring, alongside 48.7% under strict scoring reported in the system card. The accompanying system card states that Opus 5.5 scored higher than Opus 5 on every evaluation in its capability summary, with the largest gains concentrated in agentic coding, visual reasoning, computer use, and long-horizon knowledge work. The model ships with safety classifiers covering biology and cybersecurity that can route blocked requests to earlier Claude models.
—
Usage
Past 30 DaysNot available
Not in Playground
Vision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated September 22, 2026Pricing updated September 22, 2026
Claude Opus 5.5 averages 85.5% across the six Vision Evals tasks, ranking #3 of 57 models overall.
It places in the top three for Counting, Reasoning, and Object Detection.
Its weakest relative showing is OCR, ranking #37 of 57 at 87.8%.
At $0.014 per sample it is the 51st cheapest of the 57 benchmarked models, and its average inference time of 12.8s per sample makes it the 37th fastest.
Field medians: Object Detection 54.3%, Counting 61.7%, Identification 84.4%, OCR 88.7%, Data Extraction 84.5%, Reasoning 54.8%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection (low) | 74.4% ±0.5, Mean of 3 runs, range 73.9 to 74.8 | #3 of 57 | $0.022 | 17.55s | |
| Object Detection (high) | 76.8% ±1.2, Mean of 3 runs, range 75.4 to 77.8 | #3 of 23 | $0.030 | 14.07s | |
| Counting (low) | 80.6% ±2.0, Mean of 3 runs, range 78.4 to 82.4 | #2 of 57 | $0.0081 | 10.24s | |
| Counting (high) | 82.0% ±2.0, Mean of 3 runs, range 79.7 to 83.8 | #2 of 23 | $0.0098 | 11.95s | |
| Identification (low) | 93.8% ±0.0, Mean of 3 runs, range 93.8 to 93.8 | #8 of 57 | $0.0058 | 6.22s | |
| Identification (high) | 95.8% ±1.6, Mean of 3 runs, range 93.8 to 96.9 | #6 of 23 | $0.0067 | 9.81s | |
| OCR (low) | 87.8% ±0.6, Mean of 3 runs, range 87.0 to 88.2 | #37 of 57 | $0.017 | 8.45s | |
| OCR (high) | 87.2% ±0.6, Mean of 3 runs, range 86.5 to 87.8 | #22 of 23 | $0.024 | 15.90s | |
| Data Extraction (low) | 93.5% ±0.5, Mean of 3 runs, range 92.8 to 93.8 | #7 of 57 | $0.0066 | 8.80s | |
| Data Extraction (high) | 93.5% ±0.5, Mean of 3 runs, range 92.8 to 93.8 | #5 of 23 | $0.0075 | 11.54s | |
| Reasoning (low) | 83.0% ±1.0, Mean of 3 runs, range 82.1 to 84.1 | #2 of 57 | $0.0090 | 11.05s | |
| Reasoning (high) | 85.9% ±2.6, Mean of 3 runs, range 82.8 to 88.1 | #2 of 43 | $0.011 | 9.12s |
Overall benchmark score against estimated cost per sample, on a log scale. Upper-left is the sweet spot: high quality at low cost.
56 models on the current benchmark · scores and efficiency pooled across all six tasks at low effort · Claude Opus 5.5 highlighted
Claude Opus 5.5 scores are the mean of 3 runs per task at both low and high effort · Methodology
View all Vision Evals →Claude Opus 5.5 costs $4.00 per 1M input tokens and $20.00 per 1M output tokens.
Pricing updated Sep 22, 2026
Other versions in the same family as Claude Opus 5.5.
Claude Opus 5.5 is proprietary: the weights are not distributed, and the Claude Opus 5.5 license is the vendor's commercial terms of service that you accept when you call the API.
Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.
Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Claude Opus 5.5.
Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.
Talk to salesThis model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. Claude Opus 5.5 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Counting at 80.6% (#2 of 57 at low effort).
Yes. its transcriptions match the ground truth 87.8% on average (#37 of 57 at low effort) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 93.5%.
Yes. On Vision Evals, Claude Opus 5.5 scores 74.4% mAP@50 on object detection (#3 of 57 at low effort) and 80.6% judge-graded accuracy on object counting.
On our benchmark's task mix, Claude Opus 5.5 averages $0.01 per sample at $4.00 per 1M input and $20.00 per 1M output tokens (#51 of 57 on cost), with an average speed of 12.8s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Claude Opus 5.5 sits #3 of 57 at 85.5%, just behind Gemini 3.5 Flash (86%) and just ahead of Gemini 3.7 Flash (85.2%). See the full side-by-side: Claude Opus 5.5 vs Gemini 3.5 Flash.