Roboflow

This model is deprecated

GPT-4.1 and can no longer be run here. Its evaluation results and details remain available for reference. Try GPT-5.6 Sol instead.

GPT-4.1 Overview

GPT-4.1, released by OpenAI in April 2025, is a multimodal large language model that advances the GPT-4 series with major improvements in coding, reasoning, and instruction following. It accepts both text and images, supports tool calling and structured outputs, and features an expanded context window of up to ~1 million tokens—enabling it to process very large documents, multi-file codebases, or long conversations in a single prompt. Its knowledge is current through June 2024.

The GPT-4.1 family includes standard, mini, and nano variants, offering trade-offs between performance, cost, and latency. While parameter counts remain undisclosed, the series improves efficiency and responsiveness compared to GPT-4, making it suitable for both enterprise-scale tasks and cost-sensitive applications. Common use cases include software development, technical research, knowledge management, multimodal analysis, and high-context enterprise assistants.

GPT-4.1 Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Arena Rankings

GPT-4.1 Vision Evals

GPT-4.1 has been deprecated by its provider and can no longer be evaluated on the current benchmark. The legacy Vision Evals results below are preserved for reference. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#19 of 7770.15% pass rate · better than 69%
Score70.15%pass rate across 67 tasks
Speed2.56savg response per task
Cost$0.0018 / task$2.00 in · $8.00 out / 1M
Tokens977 / task891 in · 6 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Spatial Understanding16 / 19
84.2%
Defect Detection12 / 15
80%
Object Understanding11 / 14
78.6%
Document Understanding6 / 9
66.7%
Object Counting2 / 10
20%
HighestLowest
This model#18 of 5881.22% pass rate · better than 67%
Score81.22%pass rate across 229 tasks
Speed1.74savg response per task
Cost$0.0007 / task$2.00 in · $8.00 out / 1M
Tokens304 / task292 in · 10 out
Score key:≥75%40–74%<40%
CategoryPassedScore
License Plate Recognition28 / 30
93.3%
Text Recognition26 / 30
86.7%
Focused Scene OCR80 / 99
80.8%
VQA & Extraction47 / 60
78.3%
Handwritten Math5 / 10
50%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

GPT-4.1 Pricing

GPT-4.1 costs $2.00 per 1M input tokens and $8.00 per 1M output tokens.

Input$2.00 / 1M tokens
Output$8.00 / 1M tokens
Cached input$0.500 / 1M tokens

Pricing updated Aug 7, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

10 of 11 models plotted · 1 not yet evaluated

ModelScoreMedian tokensEst. cost / taskCompare
OpenAIGPT-5.6 Sol76.1%1.5K$0.0073Compare
GoogleGemini 3.1 Pro75.8%1.1K$0.0024Compare
GoogleGemini 3 Flash74.6%1.4K$0.0014Compare
OpenAIGPT-5 Mini73.1%1.8K$0.0006Compare
QwenQwen3.5 27B71.6%1.2K$0.0002Compare
OpenAIGPT-4.1(this model)70.2%977
AnthropicClaude Sonnet 570.2%2.2K$0.0048Compare
AnthropicClaude Sonnet 4.670.2%2.3K$0.0080Compare
OpenAIGPT-5.6 Luna70.2%1.5K$0.0002Compare
GoogleGemini 2.5 Pro70.2%856$0.0060Compare
GoogleGemini 3.1 Flash-Lite68.7%1.1K$0.0003Compare

Alternatives to GPT-4.1

Other models worth comparing for similar use cases.

Google
Gemini 2.5 Flash
Gemini 2.5 Flash, released on June 17, 2025, is Google DeepMind’s production-ready, efficiency-focused model in the Gemini 2.5 family. It is multimodal, accepting text, images, video, and audio as inputs, with text as the primary output format. The model supports 1 million input tokens and up to 65K output tokens, enabling it to process very large contexts such as books, long video transcripts, or extensive datasets. Its training knowledge extends to January 2025.Designed as a price-performance leader, Gemini 2.5 Flash balances speed and reasoning power, making it suitable for everyday enterprise and developer use cases without the higher latency and cost of Pro models. It supports advanced workflows like function calling, code execution, search grounding, URL context ingestion, and structured outputs. While efficient and scalable, output length is still limited compared to its input capacity, and multimodal outputs (e.g. image or audio generation) remain restricted to specialized or preview variants.
Anthropic
Claude Sonnet 4.5
Claude Sonnet 4.5, released by Anthropic in September 2025, is the company’s most advanced Sonnet-series model, built for high-performance reasoning, coding, and long-horizon agentic workflows. It is a multimodal system that accepts both text and images, with a 200,000-token context window designed for handling large documents and extended interactions. Anthropic highlights its improvements in reliability, reduced sycophancy, and alignment, making it suitable for sustained enterprise use.The model delivers strong results in coding and autonomous workflows, achieving 61.4% on the OSWorld benchmark and leading performance on SWE-bench Verified. It introduces infrastructure features such as a memory tool (beta), checkpointing for Claude Code, parallel tool use, and tighter integration with VS Code. Compared to Opus, which targets broader reasoning, Sonnet 4.5 is optimized for structured, long-duration tasks. Positioned against leading offerings from OpenAI and Google, it is aimed at enterprise automation, software engineering, and research-intensive applications.
Qwen
Qwen3 VL 235B A22B Instruct
Qwen3 VL 235B A22B Instruct is a flagship multimodal vision-language model developed by Qwen (Alibaba Cloud), designed for instruction-following tasks that combine advanced text generation with visual understanding. It serves as a high-end open-weight model for developers and researchers building multimodal AI systems that require strong reasoning, perception, and long-context capabilities.The model supports interleaved text and image inputs, very long context windows (up to roughly 256K tokens), and efficient inference through a mixture-of-experts architecture with about 22B active parameters out of 235B total. In today’s landscape, it competes with top-tier proprietary vision-language models while offering the advantages of open weights and flexible deployment. Typical applications include multimodal assistants, document and image analysis, visual reasoning, and large-context instruction-based workflows.
Google
Gemini 2.5 Pro
Gemini 2.5 Pro, released on June 17, 2025, is Google DeepMind’s most capable model in the Gemini 2.5 family, optimized for deep reasoning, coding, and complex multimodal tasks. It accepts text, images, audio, video, and PDFs as input and outputs text. The model supports 1 million input tokens with an output capacity of up to 65K tokens, enabling large-scale comprehension of datasets, codebases, and technical documents. Its training knowledge extends to January 2025.Pro outperforms earlier Gemini 2.0 models across benchmarks, including agentic coding tasks where it achieved ~63.8% on SWE-Bench Verified. It supports structured outputs, function calling, code execution, search grounding, and URL context, making it well-suited for enterprise, STEM, and developer workflows. However, it does not currently support image or audio generation in its stable release, and its higher computational cost and latency make it less efficient than Flash or Flash-Lite. It is available via the Gemini API, Google AI Studio, and Vertex AI.

Other OpenAI GPT models

Other versions in the same family as GPT-4.1.

GPT-4.1 License

Proprietary

License terms and commercial-use guidance for GPT-4.1.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About GPT-4.1 Vision

Yes. GPT-4.1 accepts image input, and on Roboflow's previous vision benchmark it passed 70.2% of visual understanding tasks (#19 of 77) and scored 81.2% on OCR.

GPT-4.1 has been deprecated by its provider and can no longer be run, so it is not part of Roboflow's current Vision Evals. Its results from the previous benchmark are preserved on this page for reference.