GPT-5.4 is a proprietary multimodal large language model developed by OpenAI and released on March 5, 2026. It is designed for professional workloads such as advanced software development, research, and agentic automation. The model combines the general reasoning capabilities of the GPT-5 series with software engineering improvements derived from GPT-5.3-Codex. In the API and Codex environments it supports context windows of up to 1 million tokens, enabling long-context reasoning and large-scale code or document workflows.
Compared with GPT-5.2, GPT-5.4 reduces false individual claims by 33% and lowers overall response errors by 18%, improving factual reliability across complex tasks. It is also the first general-purpose OpenAI release with native computer-use capabilities, allowing agents to interact with desktops, browsers, and external applications to complete multi-step workflows. The model family includes three variants: GPT-5.4 (standard), GPT-5.4 Pro for higher-performance workloads, and GPT-5.4 Thinking, a reasoning-oriented version in ChatGPT that presents an upfront plan before generating its response. The API also introduces a Tool Search system that allows models to retrieve tool definitions dynamically, reducing token usage in tool-heavy integrations.
Drag and drop an image here, or click to browse
OCR will run automatically
—
Usage
Past 30 DaysGPT-5.4 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals
| Category | Passed | Score |
|---|---|---|
| Document Understanding | 8 / 9 | 88.9% |
| Defect Detection | 13 / 15 | 86.7% |
| Object Understanding | 12 / 14 | 85.7% |
| Spatial Understanding | 15 / 19 | 78.9% |
| Object Counting | 4 / 10 | 40% |
| Category | Passed | Score |
|---|---|---|
| License Plate Recognition | 27 / 30 | 90% |
| Text Recognition | 25 / 30 | 83.3% |
| VQA & Extraction | 49 / 60 | 81.7% |
| Focused Scene OCR | 75 / 99 | 75.8% |
| Handwritten Math | 6 / 10 | 60% |
Scores based on a single evaluation run · Methodology
View all legacy Vision Evals results →GPT-5.4 costs $2.50 per 1M input tokens and $15.00 per 1M output tokens.
Pricing updated Jul 21, 2026
Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.
9 of 9 models plotted
| Model | Score | Median tokens | Est. cost / task | Compare |
|---|---|---|---|---|
| Gemini 3.5 Flash | 79.1% | 1.4K | $0.0043 | Compare |
| Claude Fable 5 | 79.1% | 2.9K | $0.041 | Compare |
| GPT-5.4 Mini | 77.6% | 1.9K | $0.0015 | Compare |
| GPT-5.4(this model) | 77.6% | 1.7K | $0.0052 | — |
| GPT-5.5 | 77.6% | 1.7K | $0.011 | Compare |
| Qwen3.5 122B A10B | 76.1% | 1.2K | $0.0003 | Compare |
| GPT-5.6 Terra | 76.1% | 1.5K | $0.0041 | Compare |
| GPT-5.6 Sol | 76.1% | 1.5K | $0.0073 | Compare |
| Gemini 3.1 Pro | 75.8% | 1.1K | $0.0024 | Compare |
Other models worth comparing for similar use cases.
Other versions in the same family as GPT-5.4.
License terms and commercial-use guidance for GPT-5.4.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. GPT-5.4 accepts image input, and on Roboflow's previous vision benchmark it passed 77.6% of visual understanding tasks (#4 of 77) and scored 79.5% on OCR. You can test it on your own image in the demo above.
GPT-5.4 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.
Yes. The demo on this page runs GPT-5.4 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.