Claude Sonnet 5 is a mid-tier large language model from Anthropic, released on June 30, 2026, as the latest model in the Sonnet series and a direct successor to Claude Sonnet 4.6. It is a hybrid reasoning model designed primarily for agentic workflows, software coding, and professional tasks. The model features a 1 million token context window, a 128k maximum output token limit, and runs adaptive thinking by default, giving API users fine-grained control over reasoning effort across five levels (low, medium, high, max, and extra-high). It uses an updated tokenizer shared with Opus 4.7 and later models, which produces approximately 30% more tokens for equivalent text compared to earlier Claude models. On benchmarks, Sonnet 5 scores 63.2% on agentic coding and 81.2% on OSWorld, narrowing the gap with Opus 4.8 while remaining at Sonnet-tier pricing.
The model supports text and image input with text output, and accepts tools including browsers and terminals for autonomous multi-step task execution. Anthropic's safety evaluations report that Sonnet 5 shows a lower rate of undesirable behaviors than Sonnet 4.6 and is generally safer in agentic contexts, with improved resistance to prompt injection and reduced sycophancy. Cybersecurity safeguards equivalent to those on Opus 4.7 and 4.8 are active, though Anthropic notes the model was not deliberately trained on cybersecurity tasks. The model is proprietary and API-only, with no open weights.
Drag and drop an image here, or click to browse
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated July 10, 2026Pricing updated July 21, 2026
Claude Sonnet 5 averages 65.6% across the six Vision Evals tasks, ranking #13 of 16 models overall.
Its weakest relative showing is Object Detection, ranking #14 of 16 at 18.0%.
At $0.0053 per sample it is the 9th cheapest of the 16 benchmarked models, and its average inference time of 4.0s per sample makes it the 2nd fastest.
Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 18.0% | #14 of 16 | $0.0071 | 4.7s | |
| Counting | 56.8% | #10 of 16 | $0.0030 | 3.0s | |
| Identification | 81.3% | #10 of 16 | $0.0027 | 2.3s | |
| OCR | 91.7% | #4 of 16 | $0.0078 | 7.2s | |
| Data Extraction | 89.7% | #5 of 16 | $0.0030 | 2.8s | |
| Reasoning | 56.5% | #12 of 16 | $0.0031 | 2.4s |
Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
16 models on the current benchmark · scores and efficiency pooled across all six tasks · Claude Sonnet 5 highlighted
Claude Sonnet 5 scores from a single evaluation run · Methodology
View all Vision Evals →Claude Sonnet 5 costs $2.00 per 1M input tokens and $10.00 per 1M output tokens.
Pricing updated Jul 21, 2026
Other models worth comparing for similar use cases.
Other versions in the same family as Claude Sonnet 5.
License terms and commercial-use guidance for Claude Sonnet 5.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. Claude Sonnet 5 accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is OCR at 91.7% (#4 of 16). You can test it on your own image in the demo above.
Yes, and it is one of the model's strongest vision skills: its transcriptions match the ground truth 91.7% on average (#4 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 89.7%.
Not its strength. On Vision Evals, Claude Sonnet 5 scores 18% mAP@50 on object detection (#14 of 16) and 56.8% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, Claude Sonnet 5 averages $0.0053 per sample at $2.00 per 1M input and $10.00 per 1M output tokens (#9 of 16 on cost), with an average speed of 4.0s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, Claude Sonnet 5 sits #13 of 16 at 65.6%, just behind GLM 5V Turbo (65.9%) and just ahead of GPT-5.4 mini (65.3%). See the full side-by-side: Claude Sonnet 5 vs GPT-5.4 mini.