Roboflow

GPT-5.4 Overview

GPT-5.4 is a proprietary multimodal large language model developed by OpenAI and released on March 5, 2026. It is designed for professional workloads such as advanced software development, research, and agentic automation. The model combines the general reasoning capabilities of the GPT-5 series with software engineering improvements derived from GPT-5.3-Codex. In the API and Codex environments it supports context windows of up to 1 million tokens, enabling long-context reasoning and large-scale code or document workflows.

Compared with GPT-5.2, GPT-5.4 reduces false individual claims by 33% and lowers overall response errors by 18%, improving factual reliability across complex tasks. It is also the first general-purpose OpenAI release with native computer-use capabilities, allowing agents to interact with desktops, browsers, and external applications to complete multi-step workflows. The model family includes three variants: GPT-5.4 (standard), GPT-5.4 Pro for higher-performance workloads, and GPT-5.4 Thinking, a reasoning-oriented version in ChatGPT that presents an upfront plan before generating its response. The API also introduces a Tool Search system that allows models to retrieve tool definitions dynamically, reducing token usage in tool-heavy integrations.

GPT-5.4 Interactive Demo

GPT-5.4 Details & Performance

Details

Resources

Vision Tasks

Vision LanguageObject DetectionClassificationOCRVisual Question AnsweringCaptioning

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Performance

Avg. Latency

Arena Rankings

GPT-5.4 Vision Evals

GPT-5.4 has not yet been evaluated on the current benchmark. The results below are from the legacy version of Vision Evals, our previous benchmark. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#4 of 7777.61% pass rate · better than 90%
Score77.61%pass rate across 67 tasks
Speed7.16savg response per task
Cost$0.0052 / task$2.50 in · $15.00 out / 1M
Tokens1.7K / task1.4K in · 108 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Document Understanding8 / 9
88.9%
Defect Detection13 / 15
86.7%
Object Understanding12 / 14
85.7%
Spatial Understanding15 / 19
78.9%
Object Counting4 / 10
40%
HighestLowest
This model#22 of 5879.48% pass rate · better than 62%
Score79.48%pass rate across 229 tasks
Speed3.98savg response per task
Cost$0.0017 / task$2.50 in · $15.00 out / 1M
Tokens300 / task105 in · 95 out
Score key:≥75%40–74%<40%
CategoryPassedScore
License Plate Recognition27 / 30
90%
Text Recognition25 / 30
83.3%
VQA & Extraction49 / 60
81.7%
Focused Scene OCR75 / 99
75.8%
Handwritten Math6 / 10
60%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

GPT-5.4 Pricing

GPT-5.4 costs $2.50 per 1M input tokens and $15.00 per 1M output tokens.

Input$2.50 / 1M tokens
Output$15.00 / 1M tokens
Cached input$0.250 / 1M tokens

Pricing updated Jul 21, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

9 of 9 models plotted

ModelScoreMedian tokensEst. cost / taskCompare
GoogleGemini 3.5 Flash79.1%1.4K$0.0043Compare
AnthropicClaude Fable 579.1%2.9K$0.041Compare
OpenAIGPT-5.4 Mini77.6%1.9K$0.0015Compare
OpenAIGPT-5.4(this model)77.6%1.7K$0.0052
OpenAIGPT-5.577.6%1.7K$0.011Compare
QwenQwen3.5 122B A10B76.1%1.2K$0.0003Compare
OpenAIGPT-5.6 Terra76.1%1.5K$0.0041Compare
OpenAIGPT-5.6 Sol76.1%1.5K$0.0073Compare
GoogleGemini 3.1 Pro75.8%1.1K$0.0024Compare

Alternatives to GPT-5.4

Other models worth comparing for similar use cases.

Grok
Grok 4
Grok 4, released by xAI on July 9, 2025, is the fourth-generation model in the Grok family and the most advanced to date. It is multimodal, supporting text, vision, tool use, and real-time web search, with a reported 256,000-token context window for long-form reasoning and document analysis. Its training data extends through November 2024, making it the most up-to-date Grok model at launch.The lineup includes Grok 4 Generalist for broad tasks, Grok 4 Heavy for higher-capacity reasoning, and Grok 4 Code optimized for programming and debugging. A notable feature is its always-on “Think” mode, designed for deeper multi-step reasoning. While xAI has not disclosed parameter counts, Grok 4 is positioned to compete with frontier models like GPT-5 and Claude 4, balancing real-time knowledge via web integration with structured tool use. It is best suited for coding, complex reasoning, and multimodal AI assistants.
Anthropic
Claude Opus 4.7
Claude Opus 4.7 is a proprietary multimodal language model developed by Anthropic, released on April 16, 2026. It is designed for agentic coding, long-horizon task execution, and enterprise knowledge work. The model supports text and vision inputs and operates with a context window of up to 1,000,000 tokens. It introduces adaptive thinking, which dynamically allocates reasoning based on task complexity, along with configurable effort controls including a new xhigh setting that sits between the existing high and max levels. It achieves 87.6% on SWE-bench Verified and 78.0% on OSWorld-Verified, reflecting strong performance on autonomous software engineering and computer use tasks respectively.Compared to Claude Opus 4.6, version 4.7 shows improved instruction following and higher reliability in extended agentic tasks. Vision capabilities now support high-resolution inputs up to 2,576px on the long edge (~3.75 megapixels), more than three times the resolution of prior Claude models, enabling finer interpretation of dense diagrams, UI screenshots, and document layouts. These improvements, combined with self-verification on long-running tasks and a new task budget system for controlling agentic loops, make it well-suited for complex software engineering, technical analysis, and multimodal vision workflows.
Google
Gemini 3.1 Pro
Gemini 3.1 Pro is a proprietary multimodal model from Google’s Gemini 3 series, released in early 2026 and designed for advanced reasoning across large multimodal datasets. It accepts text, images, audio, video, and documents, supporting up to a 1-million-token input context with up to 64k output tokens. Compared with Gemini 3 Pro, it improves long-context synthesis and multi-step reasoning, enabling more reliable analysis of large documents, datasets, and software codebases.The model also advances visual understanding and grounding, allowing it to interpret UI screenshots, diagrams, and real-world scenes while referencing specific regions within images or video. These capabilities make Gemini 3.1 Pro well suited for multimodal workflows involving document processing, interface analysis, robotics research, and complex visual reasoning.

Other OpenAI GPT models

Other versions in the same family as GPT-5.4.

GPT-5.4 License

Proprietary

License terms and commercial-use guidance for GPT-5.4.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About GPT-5.4 Vision

Yes. GPT-5.4 accepts image input, and on Roboflow's previous vision benchmark it passed 77.6% of visual understanding tasks (#4 of 77) and scored 79.5% on OCR. You can test it on your own image in the demo above.

GPT-5.4 has not yet been evaluated on Roboflow's current Vision Evals. The results on this page are from the previous benchmark.

Yes. The demo on this page runs GPT-5.4 in the free Roboflow Playground: upload an image and see results in seconds. A free account unlocks unlimited runs.