This model is deprecated
GPT-4.1 and can no longer be run here. Its evaluation results and details remain available for reference. Try GPT-5.6 Sol instead.
GPT-4.1, released by OpenAI in April 2025, is a multimodal large language model that advances the GPT-4 series with major improvements in coding, reasoning, and instruction following. It accepts both text and images, supports tool calling and structured outputs, and features an expanded context window of up to ~1 million tokens—enabling it to process very large documents, multi-file codebases, or long conversations in a single prompt. Its knowledge is current through June 2024.
The GPT-4.1 family includes standard, mini, and nano variants, offering trade-offs between performance, cost, and latency. While parameter counts remain undisclosed, the series improves efficiency and responsiveness compared to GPT-4, making it suitable for both enterprise-scale tasks and cost-sensitive applications. Common use cases include software development, technical research, knowledge management, multimodal analysis, and high-context enterprise assistants.
—
Usage
Past 30 DaysNot available
Not in Playground
GPT-4.1 has been deprecated by its provider and can no longer be evaluated on the current benchmark. The legacy Vision Evals results below are preserved for reference. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Spatial Understanding | 16 / 19 | 84.2% |
| Defect Detection | 12 / 15 | 80% |
| Object Understanding | 11 / 14 | 78.6% |
| Document Understanding | 6 / 9 | 66.7% |
| Object Counting | 2 / 10 | 20% |
| Category | Passed | Score |
|---|---|---|
| License Plate Recognition | 28 / 30 | 93.3% |
| Text Recognition | 26 / 30 | 86.7% |
| Focused Scene OCR | 80 / 99 | 80.8% |
| VQA & Extraction | 47 / 60 | 78.3% |
| Handwritten Math | 5 / 10 | 50% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →GPT-4.1 costs $2.00 per 1M input tokens and $8.00 per 1M output tokens.
Pricing updated Aug 7, 2026
Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
10 of 11 models plotted · 1 not yet evaluated
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| GPT-5.6 Sol | 76.1% | 1.5K | $0.0073 | Compare |
| Gemini 3.1 Pro | 75.8% | 1.1K | $0.0024 | Compare |
| Gemini 3 Flash | 74.6% | 1.4K | $0.0014 | Compare |
| GPT-5 Mini | 73.1% | 1.8K | $0.0006 | Compare |
| Qwen3.5 27B | 71.6% | 1.2K | $0.0002 | Compare |
| GPT-4.1(this model) | 70.2% | 977 | — | — |
| Claude Sonnet 5 | 70.2% | 2.2K | $0.0048 | Compare |
| Claude Sonnet 4.6 | 70.2% | 2.3K | $0.0080 | Compare |
| GPT-5.6 Luna | 70.2% | 1.5K | $0.0002 | Compare |
| Gemini 2.5 Pro | 70.2% | 856 | $0.0060 | Compare |
| Gemini 3.1 Flash-Lite | 68.7% | 1.1K | $0.0003 | Compare |
Other models worth comparing for similar use cases.
License terms and commercial-use guidance for GPT-4.1.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. GPT-4.1 accepts image input, and on Roboflow's previous vision benchmark it passed 70.2% of visual understanding tasks (#19 of 77) and scored 81.2% on OCR.
GPT-4.1 has been deprecated by its provider and can no longer be run, so it is not part of Roboflow's current Vision Evals. Its results from the previous benchmark are preserved on this page for reference.