Roboflow
Anthropic

Anthropic: Claude Sonnet 4

This model is deprecated

Claude Sonnet 4 was deprecated on Jun 15, 2026 and can no longer be run here. Its evaluation results and details remain available for reference. Try Claude Sonnet 5 instead.

Claude Sonnet 4 Overview

Claude 4 Sonnet, released by Anthropic in May 2025, is the mid-tier model in the Claude 4 family, designed to balance capability, cost, and speed. It is multimodal, accepting both text and images, and extends beyond prior versions with improved “computer use” support, allowing API-driven interaction with desktop-like interfaces. By default, it supports 200,000 tokens of context, but as of August 2025, it also offers a 1 million-token context window in public beta—making it one of the most context-capable models available for processing entire codebases or large document sets in a single request.

Sonnet 4 is significantly cheaper than the flagship Opus while still demonstrating strong reasoning, coding, and instruction-following ability with reduced hallucinations. Its extended context capabilities and lower latency make it well-suited for enterprise-scale knowledge management, software development, research assistants, and productivity automation where both cost efficiency and high reliability are essential.

Claude Sonnet 4 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Arena Rankings

Claude Sonnet 4 Vision Evals

Claude Sonnet 4 has been deprecated by its provider and can no longer be evaluated on the current benchmark. The legacy Vision Evals results below are preserved for reference. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#25 of 7768.66% pass rate · better than 62%
Score68.66%pass rate across 67 tasks
Speed21.26savg response per task
Cost$3.00 in · $15.00 out / 1M
Tokenstokens unavailable
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding8 / 9
88.9%
Defect Detection12 / 15
80%
Object Understanding11 / 14
78.6%
Spatial Understanding13 / 19
68.4%
Object Counting2 / 10
20%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Claude Sonnet 4 Pricing

Claude Sonnet 4 costs $3.00 per 1M input tokens and $15.00 per 1M output tokens.

Input$3.00 / 1M tokens
Output$15.00 / 1M tokens
Cached input$0.300 / 1M tokens

Pricing updated Aug 5, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

10 of 11 models plotted · 1 not yet evaluated

Alternatives to Claude Sonnet 4

Other models worth comparing for similar use cases.

OpenAI
GPT-5 Mini
GPT-5 Mini, released by OpenAI on August 7, 2025, is a mid-tier variant of the GPT-5 family that balances cost, speed, and capability. It is multimodal, supporting both text and image inputs, and offers a substantial input context window of ~400,000 tokens with output lengths up to ~128,000 tokens. While less powerful than the full GPT-5, it inherits its safety tuning, instruction-following improvements, and multimodal reasoning, making it a practical choice for developers who need large context handling without the expense of premium models.GPT-5 Mini is optimized for affordability while retaining strong reasoning performance. Benchmarks show it outperforming earlier models such as GPT-4o on many multimodal and medical VQA tasks, though it lags behind GPT-5 on the most complex problems. Ideal use cases include prototyping, scalable content generation, document analysis, and mid-range reasoning tasks where efficiency and context capacity matter more than top-tier accuracy.
Google
Gemini 2.5 Flash
Gemini 2.5 Flash, released on June 17, 2025, is Google DeepMind’s production-ready, efficiency-focused model in the Gemini 2.5 family. It is multimodal, accepting text, images, video, and audio as inputs, with text as the primary output format. The model supports 1 million input tokens and up to 65K output tokens, enabling it to process very large contexts such as books, long video transcripts, or extensive datasets. Its training knowledge extends to January 2025.Designed as a price-performance leader, Gemini 2.5 Flash balances speed and reasoning power, making it suitable for everyday enterprise and developer use cases without the higher latency and cost of Pro models. It supports advanced workflows like function calling, code execution, search grounding, URL context ingestion, and structured outputs. While efficient and scalable, output length is still limited compared to its input capacity, and multimodal outputs (e.g. image or audio generation) remain restricted to specialized or preview variants.
Google
Gemini 3.5 Flash
Gemini 3.5 Flash is a multimodal language model developed by Google DeepMind and released at Google I/O 2026. It is built on the Gemini 3 Flash reasoning foundation and introduces configurable thinking levels (minimal, low, medium, and high) that allow developers to tune the depth of internal reasoning before a response is generated. The model accepts text, image, video, audio, and PDF inputs and produces text output, with a 1 million token context window and up to 65,000 output tokens per request. It is natively multimodal, processing visual inputs alongside text to support tasks such as image captioning, classification, optical character recognition, object detection, and visual grounding, where the model references specific regions within an image or video frame.Its vision capabilities extend to interpreting UI screenshots, diagrams, charts, and real-world scenes, as well as understanding video and live frame sequences for activity and scene recognition. The model supports combined tool use, including Google Search, URL context, code execution, and custom functions, within a single request, and it uses reasoning context from previous turns when thought signatures are present in the conversation history, enabling persistent multi-turn reasoning chains. Gemini 3.5 Flash carries a knowledge cutoff of January 2026 and is available via the Gemini API, Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform.

Other Anthropic Sonnet models

Other versions in the same family as Claude Sonnet 4.

Claude Sonnet 4 License

Proprietary

License terms and commercial-use guidance for Claude Sonnet 4.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Claude Sonnet 4 Vision

Yes. Claude Sonnet 4 accepts image input, and on Roboflow's previous vision benchmark it passed 68.7% of visual understanding tasks (#25 of 77).

Claude Sonnet 4 has been deprecated by its provider and can no longer be run, so it is not part of Roboflow's current Vision Evals. Its results from the previous benchmark are preserved on this page for reference.