Roboflow
Anthropic

Anthropic: Claude Opus 4.7

Claude Opus 4.7 Overview

Claude Opus 4.7 is a proprietary multimodal language model developed by Anthropic, released on April 16, 2026. It is designed for agentic coding, long-horizon task execution, and enterprise knowledge work. The model supports text and vision inputs and operates with a context window of up to 1,000,000 tokens. It introduces adaptive thinking, which dynamically allocates reasoning based on task complexity, along with configurable effort controls including a new xhigh setting that sits between the existing high and max levels. It achieves 87.6% on SWE-bench Verified and 78.0% on OSWorld-Verified, reflecting strong performance on autonomous software engineering and computer use tasks respectively.

Compared to Claude Opus 4.6, version 4.7 shows improved instruction following and higher reliability in extended agentic tasks. Vision capabilities now support high-resolution inputs up to 2,576px on the long edge (~3.75 megapixels), more than three times the resolution of prior Claude models, enabling finer interpretation of dense diagrams, UI screenshots, and document layouts. These improvements, combined with self-verification on long-running tasks and a new task budget system for controlling agentic loops, make it well-suited for complex software engineering, technical analysis, and multimodal vision workflows.

Claude Opus 4.7 Interactive Demo

Claude Opus 4.7 Details & Performance

Details

Resources

Vision Tasks

Object DetectionClassificationOCRVision LanguageCaptioningVisual Question Answering

Features

Foundation VisionMultimodal VisionLLMs with Vision Capabilities

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

Claude Opus 4.7 Vision Evals

Claude Opus 4.7 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#30 of 7767.16% pass rate · better than 55%
Score67.16%pass rate across 67 tasks
Speed4.85savg response per task
Cost$0.015 / task$5.00 in · $25.00 out / 1M
Tokens2.6K / task2.4K in · 110 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Object Understanding12 / 14
85.7%
Document Understanding7 / 9
77.8%
Defect Detection11 / 15
73.3%
Spatial Understanding13 / 19
68.4%
Object Counting2 / 10
20%
HighestLowest
This model#10 of 5886.9% pass rate · better than 83%
Score86.9%pass rate across 229 tasks
Speed4.19savg response per task
Cost$0.0069 / task$5.00 in · $25.00 out / 1M
Tokens1.1K / task969 in · 81 out
Score key:≥75%40–74%<40%
CategoryPassedScore
License Plate Recognition28 / 30
93.3%
Focused Scene OCR88 / 99
88.9%
Text Recognition26 / 30
86.7%
VQA & Extraction49 / 60
81.7%
Handwritten Math8 / 10
80%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

Claude Opus 4.7 Pricing

Claude Opus 4.7 costs $5.00 per 1M input tokens and $25.00 per 1M output tokens.

Input$5.00 / 1M tokens
Output$25.00 / 1M tokens
Cached input$0.500 / 1M tokens

Pricing updated Jul 21, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

11 of 11 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
GoogleGemini 2.5 Pro70.2%856$0.0060Compare
GoogleGemini 3.1 Flash-Lite68.7%1.1K$0.0003Compare
GoogleGemma 4 26B A4B68.7%531$0.0001Compare
QwenQwen3.6 Plus68.7%1.6K$0.0005Compare
AnthropicClaude Opus 4.867.2%2.2K$0.012Compare
AnthropicClaude Opus 4.7(this model)67.2%2.6K$0.015
GoogleGemma 4 31B67.2%467$0.0001Compare
AnthropicClaude Opus 4.6 64.2%2.3K$0.014Compare
OpenAIGPT-5.4 Nano62.7%1.8K$0.0004Compare
MetaLlama 4 Maverick59.7%2.4K$0.0005Compare
AnthropicClaude Sonnet 4.559.7%2.3K$0.0092Compare

Alternatives to Claude Opus 4.7

Other models worth comparing for similar use cases.

Anthropic
Claude Fable 5
Claude Fable 5 is Anthropic's first generally available Mythos-class large language model, released on June 9, 2026. It is built for long-horizon, asynchronous, and agentic tasks that prior Claude generations could not sustain, including multi-day autonomous coding sessions, complex knowledge work, and document-heavy analysis. The model supports a 1 million token context window with up to 128,000 output tokens per request and uses adaptive thinking as its sole reasoning mode, where the effort level is adjustable but raw chain-of-thought is never returned. Vision capabilities allow the model to parse diagrams, charts, and tables embedded in files and PDFs, and to use visual feedback to evaluate its own coding outputs against design goals. On benchmarks such as SWE-Bench Pro, the model scores 80.3% compared to 69.2% for Claude Opus 4.8, and it leads on CursorBench 3.1 for autonomous coding workflows.Claude Fable 5 shares the same underlying model weights as Claude Mythos 5, but is deployed with safety classifiers that automatically reroute queries in high-risk domains — including cybersecurity, biology, and chemistry — to Claude Opus 4.8. These classifiers trigger in fewer than 5% of sessions on average. As a designated Covered Model, all traffic is subject to mandatory 30-day data retention to support safety monitoring. The model is available via the Claude API, Amazon Bedrock, Vertex AI, and Microsoft Foundry. Anthropic has not publicly disclosed parameter count, architecture details, or training data composition for this model.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.
OpenAI
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 family, which also includes Terra (a balanced everyday-work tier) and Luna (a fast, cost-efficient tier). Sol is designed for demanding reasoning, long-horizon agentic workflows, software engineering, computer use, scientific research, and cybersecurity tasks. It introduces two new capability modes: a "max" reasoning effort setting that allocates additional compute time for difficult problems, and an "ultra" mode that coordinates multiple subagents in parallel to accelerate complex, multi-step work. The model supports native multimodal input, allowing it to process screenshots, diagrams, charts, documents, and photographs alongside text. A reported context window of approximately 1.5 million tokens enables processing of large codebases, lengthy research documents, and extended agentic sessions.GPT-5.6 Sol was announced on June 26, 2026, initially in a limited preview for trusted partners, and reached general availability on July 9, 2026. On the Agents' Last Exam benchmark, which evaluates long-running professional workflows across 55 fields, Sol scores 53.6. On Terminal-Bench 2.1, which tests command-line agentic coding workflows, Sol Ultra achieves 91.9%. The model also demonstrates gains in life sciences evaluations, including long-horizon genomics and quantitative biology analyses. OpenAI paired the release with its most extensive safety evaluation to date, combining human red teaming with large-scale automated testing, and classified Sol as High capability in both cybersecurity and biological risk under its Preparedness Framework, though it does not cross the Critical threshold in either category.
Qwen
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.
MoonshotAI
Kimi K3
Kimi K3 is a sparse Mixture-of-Experts large language model developed by Moonshot AI, with 2.8 trillion total parameters and a 1-million-token context window. The model activates 16 out of 896 experts per token using the Stable LatentMoE framework, and is built on two architectural innovations: Kimi Delta Attention (KDA), a hybrid linear attention mechanism that enables up to 6.3x faster decoding in long-context settings, and Attention Residuals (AttnRes), which selectively retrieves representations across model depth and delivers roughly 25% higher training efficiency. Together with refined training and data recipes, these structural advances yield approximately 2.5x better overall scaling efficiency compared to its predecessor Kimi K2. The model applies quantization-aware training from the supervised fine-tuning stage onward, using MXFP4 weights with MXFP8 activations for hardware compatibility. Thinking mode is always enabled at launch, with reasoning effort configurable via the reasoning_effort field.Kimi K3 supports native visual understanding alongside text, accepting image inputs for tasks that combine software engineering and visual reasoning. It targets long-horizon coding, knowledge work, and agentic workflows, and ships in two variants: K3 Max for general chat and agent tasks, and K3 Swarm Max for large-scale parallel processing across many coordinated sub-agents. The model is compatible with the OpenAI SDK via an OpenAI-compatible API. Full model weights are scheduled for release by July 27, 2026 under a Modified MIT license, following the open-weight pattern established by the Kimi K2 model family. A technical report with full architecture, training, and evaluation details is expected to accompany the weights release.

Other Anthropic Opus models

Other versions in the same family as Claude Opus 4.7.

Claude Opus 4.7 License

Proprietary

License terms and commercial-use guidance for Claude Opus 4.7.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About Claude Opus 4.7 Vision

Yes. Claude Opus 4.7 accepts image input, and on Roboflow's previous vision benchmark it passed 67.2% of visual understanding tasks (#30 of 77) and scored 86.9% on OCR. You can test it on your own image in the demo above.

Claude Opus 4.7 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs Claude Opus 4.7 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.