Roboflow
Anthropic

Anthropic: Claude Sonnet 4

This model is deprecated

Claude Sonnet 4 was deprecated on Jun 15, 2026 and can no longer be run here. Its evaluation results and details remain available for reference. Try Claude Sonnet 5 instead.

Claude Sonnet 4 Overview

Claude 4 Sonnet, released by Anthropic in May 2025, is the mid-tier model in the Claude 4 family, designed to balance capability, cost, and speed. It is multimodal, accepting both text and images, and extends beyond prior versions with improved “computer use” support, allowing API-driven interaction with desktop-like interfaces. By default, it supports 200,000 tokens of context, but as of August 2025, it also offers a 1 million-token context window in public beta—making it one of the most context-capable models available for processing entire codebases or large document sets in a single request.

Sonnet 4 is significantly cheaper than the flagship Opus while still demonstrating strong reasoning, coding, and instruction-following ability with reduced hallucinations. Its extended context capabilities and lower latency make it well-suited for enterprise-scale knowledge management, software development, research assistants, and productivity automation where both cost efficiency and high reliability are essential.

Claude Sonnet 4 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Claude Sonnet 4 Vision Evals

Claude Sonnet 4 has been deprecated by its provider and can no longer be evaluated on the current benchmark. The legacy Vision Evals results below are preserved for reference. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#25 of 7768.66% pass rate · better than 62%
Score68.66%pass rate across 67 tasks
Speed21.26savg response per task
Cost$3.00 in · $15.00 out / 1M
Tokenstokens unavailable
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding8 / 9
88.9%
Defect Detection12 / 15
80%
Object Understanding11 / 14
78.6%
Spatial Understanding13 / 19
68.4%
Object Counting2 / 10
20%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Claude Sonnet 4 Pricing

Claude Sonnet 4 costs $3.00 per 1M input tokens and $15.00 per 1M output tokens.

Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
Cached input$0.300 / 1M tokens

Pricing updated Sep 18, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

10 of 11 models plotted · 1 not yet evaluated

Alternatives to Claude Sonnet 4

Other models worth comparing for similar use cases.

OpenAI
GPT-5 Mini
GPT-5 Mini, released by OpenAI on August 7, 2025, is a mid-tier variant of the GPT-5 family that balances cost, speed, and capability. It is multimodal, supporting both text and image inputs, and offers a substantial input context window of ~400,000 tokens with output lengths up to ~128,000 tokens. While less powerful than the full GPT-5, it inherits its safety tuning, instruction-following improvements, and multimodal reasoning, making it a practical choice for developers who need large context handling without the expense of premium models.GPT-5 Mini is optimized for affordability while retaining strong reasoning performance. Benchmarks show it outperforming earlier models such as GPT-4o on many multimodal and medical VQA tasks, though it lags behind GPT-5 on the most complex problems. Ideal use cases include prototyping, scalable content generation, document analysis, and mid-range reasoning tasks where efficiency and context capacity matter more than top-tier accuracy.
Google
Gemini 2.5 Flash
Gemini 2.5 Flash, released on June 17, 2025, is Google DeepMind’s production-ready, efficiency-focused model in the Gemini 2.5 family. It is multimodal, accepting text, images, video, and audio as inputs, with text as the primary output format. The model supports 1 million input tokens and up to 65K output tokens, enabling it to process very large contexts such as books, long video transcripts, or extensive datasets. Its training knowledge extends to January 2025.Designed as a price-performance leader, Gemini 2.5 Flash balances speed and reasoning power, making it suitable for everyday enterprise and developer use cases without the higher latency and cost of Pro models. It supports advanced workflows like function calling, code execution, search grounding, URL context ingestion, and structured outputs. While efficient and scalable, output length is still limited compared to its input capacity, and multimodal outputs (e.g. image or audio generation) remain restricted to specialized or preview variants.
Qwen
Qwen3 VL 30B A3B Instruct
Qwen3 VL 30B A3B Instruct is an open-weight multimodal large language model developed by Alibaba as part of the Qwen family, built for instruction-following tasks that unify text generation with visual and video understanding. Released around October 2025 under the Apache-2.0 license, it targets efficient, high-fidelity vision-language reasoning across very long contexts.The model accepts text and image inputs and produces text outputs, with strong performance in OCR, spatial reasoning, long-video understanding, and agentic or GUI-centric visual tasks. It uses a Mixture-of-Experts (A3B) design with ~31.1B total parameters and ~3B active per token, paired with Qwen3-VL’s unified multimodal stack (including Interleaved-MRoPE and DeepStack fusion) to process text, images, and video in a single architecture. OCR support expands to 32 languages, enhancing document workflows. With a native ~262K token context window (extendable further), it stands out today for its balance of scale, efficiency, long-context support, and open accessibility in multimodal systems.

Other Anthropic Sonnet models

Other versions in the same family as Claude Sonnet 4.

Claude Sonnet 4 License

Proprietary

Claude Sonnet 4 is proprietary: the weights are not distributed, and the Claude Sonnet 4 license is the vendor's commercial terms of service that you accept when you call the API.

Commercial use
Permitted under the vendor terms, typically metered per token or per request, with the vendor usage policy applying to your inputs and outputs.
Modification
Not available. Claude Sonnet 4 weights are closed, so you can configure prompts and use vendor-hosted fine-tuning where it is offered, but you cannot modify the model itself.
Redistribution
Not permitted. You cannot self-host or resell the model; you build on the hosted API instead.

Vendor terms govern data retention, whether your inputs can be trained on, rate limits, and regional availability, and they can change with notice. Review them if you handle regulated or customer data.

Do I need a commercial license for Claude Sonnet 4?

Proprietary terms are set by the vendor rather than negotiated per project, and no open-source obligation attaches to your code. If you would rather deploy a model whose commercial license is included in your plan — on Roboflow Managed Cloud or a Self-Hosted Inference Server — Roboflow's licensing page lists the supported alternatives to Claude Sonnet 4.

Do not hesitate to reach out with questions for your commercial project — our team will help you start solving business problems on the first call. See Roboflow commercial licensing for the models included in each plan.

Talk to sales

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Claude Sonnet 4 Vision

Yes. Claude Sonnet 4 accepts image input, and on Roboflow's previous vision benchmark it passed 68.7% of visual understanding tasks (#25 of 77).

Claude Sonnet 4 has been deprecated by its provider and can no longer be run, so it is not part of Roboflow's current Vision Evals. Its results from the previous benchmark are preserved on this page for reference.