GPT-5.4 mini is a fast, cost-efficient model developed by OpenAI and released on March 17, 2026, optimized for high-throughput workloads and subagent orchestration. It supports text and image inputs within a 400,000-token context window, making it ideal for processing extensive visual datasets and large codebases in a single request. Designed for low-latency production environments, the model integrates with key API features including function calling, web search, and tool-based computer use, allowing it to assist in automated workflows that require navigating digital interfaces.
Compared to the previous GPT-5 mini, this version runs more than twice as fast while approaching the performance levels of the flagship GPT-5.4 on reasoning and coding benchmarks. While the larger GPT-5.4 introduces native, state-of-the-art computer-use capabilities, GPT-5.4 mini provides a scalable alternative for interpreting screenshots and reasoning over dense UI layouts. For vision tasks on Playground, it excels at extracting structured information from visual documents and assisting in agentic tasks that involve real-time interpretation of software interfaces alongside text.
Drag and drop an image here, or click to browse
—
Usage
Past 30 DaysVision Evals is Roboflow's ground-truth benchmark: every model runs the same real-world samples across six vision tasks, and answers are scored against ground truth.
Evals updated July 10, 2026Pricing updated July 21, 2026
GPT-5.4 Mini averages 65.3% across the six Vision Evals tasks, ranking #14 of 16 models overall.
Its weakest relative showing is Object Detection, ranking #16 of 16 at 3.9%.
At $0.0026 per sample it is the 5th cheapest of the 16 benchmarked models, and its average inference time of 4.5s per sample makes it the 4th fastest.
Field medians: Object Detection 41.5%, Counting 62.2%, Identification 84.4%, OCR 89.1%, Data Extraction 85.6%, Reasoning 76.1%.
| Task | Score | Field (0 to 100) | Rank | Cost / sample | Speed |
|---|---|---|---|---|---|
| Object Detection | 3.9% | #16 of 16 | $0.0032 | 5.3s | |
| Counting | 60.8% | #9 of 16 | $0.0019 | 4.0s | |
| Identification | 78.1% | #12 of 16 | $0.0013 | 2.7s | |
| OCR | 88.1% | #14 of 16 | $0.0042 | 6.2s | |
| Data Extraction | 82.5% | #11 of 16 | $0.0014 | 2.8s | |
| Reasoning | 78.3% | #7 of 16 | $0.0016 | 4.5s |
Overall benchmark score against estimated cost per sample. Upper-left is the sweet spot: high quality at low cost.
16 models on the current benchmark · scores and efficiency pooled across all six tasks · GPT-5.4 mini highlighted
GPT-5.4 Mini scores from a single evaluation run · Methodology
View all Vision Evals →GPT-5.4 Mini costs $0.750 per 1M input tokens and $4.50 per 1M output tokens.
Pricing updated Jul 21, 2026
Other models worth comparing for similar use cases.
Other versions in the same family as GPT-5.4 Mini.
License terms and commercial-use guidance for GPT-5.4 Mini.
This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.
Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.
License information is provided as a guide and is not legal advice.
Yes. GPT-5.4 Mini accepts image input and handles OCR, data extraction, object counting, identification, visual reasoning, and object detection. On Roboflow's Vision Evals its strongest task is Reasoning at 78.3% (#7 of 16). You can test it on your own image in the demo above.
Yes. its transcriptions match the ground truth 88.1% on average (#14 of 16) on Vision Evals OCR. Pulling specific fields out of documents (data extraction) scores 82.5%.
Not its strength. On Vision Evals, GPT-5.4 Mini scores 3.9% mAP@50 on object detection (#16 of 16) and 60.8% exact-match accuracy on object counting. For production counting or precise localization, pairing it with a specialized detector like RF-DETR or your own trained model in a Roboflow Workflow is usually more reliable: detect the objects, then count the detections.
On our benchmark's task mix, GPT-5.4 Mini averages $0.0026 per sample at $0.75 per 1M input and $4.50 per 1M output tokens (#5 of 16 on cost), with an average speed of 4.5s per sample across the benchmark. Actual cost depends on your images and prompts.
On the overall Vision Evals ranking, GPT-5.4 Mini sits #14 of 16 at 65.3%, just behind Claude Sonnet 5 (65.6%) and just ahead of Claude Opus 4.8 (64.8%). See the full side-by-side: GPT-5.4 Mini vs Claude Sonnet 5.