Roboflow
OpenAI

OpenAI: GPT-4.1 nano

This model is deprecated

GPT-4.1 nano and can no longer be run here. Its evaluation results and details remain available for reference. Try GPT-5.6 Luna instead.

GPT-4.1 nano Overview

GPT-4.1 nano, released by OpenAI in April 2025, is the smallest and most cost-efficient member of the GPT-4.1 family. It is multimodal, supporting both text and image inputs, and retains the family’s extended 1 million-token context window—allowing it to handle large documents or codebases despite its lightweight design. Its training knowledge extends to June 2024.

GPT-4.1 nano prioritizes speed and affordability over raw reasoning power. While less capable than GPT-4.1 and GPT-4.1 mini, it is well-suited for high-volume or latency-sensitive workloads such as classification, autocomplete, content moderation, and lightweight assistants. This makes it an attractive option for developers seeking scalable deployment where efficiency is more critical than deep reasoning.

GPT-4.1 nano Details & Performance

Details

Resources

Vision Tasks

CaptioningChart Question AnsweringClassificationDocument Question AnsweringImage TaggingMulti-Label ClassificationOCRObject DetectionVision LanguageVisual Question Answering

Features

Foundation VisionLLMs with Vision CapabilitiesMultimodal Vision

Usage

Past 30 Days

Not available

Not in Playground

Performance

Avg. Latency

Arena Rankings

GPT-4.1 nano Vision Evals

GPT-4.1 nano has been deprecated by its provider and can no longer be evaluated on the current benchmark. The legacy Vision Evals results below are preserved for reference. See the current Vision Evals

Visual Understanding

77 models · 67 tasks
HighestLowest
This model#69 of 7740.3% pass rate · better than 9%
Score40.3%pass rate across 67 tasks
Speed2.36savg response per task
Cost$0.0003 / task$0.100 in · $0.400 out / 1M
Tokens2.9K / task2.9K in · 6 out
Score key:≥75%40–74%<40%
CategoryPassedScore
Object Understanding9 / 14
64.3%
Spatial Understanding8 / 19
42.1%
Defect Detection6 / 15
40%
Document Understanding3 / 9
33.3%
Object Counting1 / 10
10%
HighestLowest
This model#52 of 5857.21% pass rate · better than 10%
Score57.21%pass rate across 229 tasks
Speed1.39savg response per task
Cost<$0.0001 / task$0.100 in · $0.400 out / 1M
Tokens189 / task175 in · 9 out
Score key:≥75%40–74%<40%
CategoryPassedScore
VQA & Extraction39 / 60
65%
Text Recognition19 / 30
63.3%
Focused Scene OCR54 / 99
54.5%
License Plate Recognition15 / 30
50%
Handwritten Math4 / 10
40%

Scores based on a single evaluation run · Methodology

View all legacy Vision Evals results →

GPT-4.1 nano Pricing

GPT-4.1 nano costs $0.100 per 1M input tokens and $0.400 per 1M output tokens.

Input$0.100 / 1M tokens
Output$0.400 / 1M tokens
Cached input$0.025 / 1M tokens

Pricing updated Aug 7, 2026

Price vs. performance

Estimated cost per task vs. Visual Understanding score, for this model and others ranked near it. Upper-left is the sweet spot (high quality, low cost). Based on Vision Evals (legacy) results.

6 of 7 models plotted · 1 not yet evaluated

ModelScoreMedian tokensEst. cost / taskCompare
AnthropicClaude Haiku 4.558.2%2.3K$0.0030Compare
OpenAIGPT-5 Nano58.2%2.7K$0.0003Compare
QwenQwen3.5 397B A17B58.2%1.5K$0.0006Compare
GoogleGemini 2.5 Flash55.2%476$0.0005Compare
GoogleGemini 2.5 Flash-Lite53.7%301<$0.0001Compare
OpenAIGPT-4.1 Nano(this model)40.3%2.9K
MoonshotAIKimi K2.535.8%2.7K$0.0031Compare

Alternatives to GPT-4.1 nano

Other models worth comparing for similar use cases.

Google
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash-Lite is a natively multimodal reasoning model from Google DeepMind in the Gemini 3 series, based on the Gemini 3 Pro architecture. It processes text, image, video, audio, and PDF inputs within a 1 million token context window and produces text output up to 64K tokens. The model targets high-volume, latency-sensitive workloads and supports visual question answering, image and document data extraction, content moderation, classification, translation, automated speech recognition, and agentic data pipelines. It exposes configurable thinking levels of minimal, low, medium, and high, which set the depth of internal reasoning applied per request and let developers balance response quality against cost and latency.On benchmarks reported at launch, Gemini 3.1 Flash-Lite scores 86.9% on GPQA Diamond and 76.8% on the MMMU Pro multimodal benchmark, and reaches an Elo score of 1432 on the Arena.ai leaderboard. According to Artificial Analysis benchmarks, it produces a 2.5 times faster time to first answer token and a 45% increase in output speed relative to Gemini 2.5 Flash. It also shows improved instruction following, higher audio input quality for automated speech recognition tasks, and support for structured JSON output used in data extraction pipelines.
Anthropic
Claude Haiku 4.5
Claude Haiku 4.5 is Anthropic’s lightweight model in the Claude 4.5 series, released in October 2025 under a proprietary license. Designed for speed and cost efficiency, it delivers near-frontier performance while maintaining Anthropic’s AI Safety Level 2 standard. Haiku 4.5 supports both text and multimodal (text and image) inputs, integrates tool use and extended reasoning, and features a 200,000 token context window, making it adept at handling long or complex workflows. Though the parameter count remains undisclosed, it achieves about 73.3% on SWE-bench Verified, reflecting strong coding and reasoning ability. Haiku 4.5 is ideal for developers and researchers seeking rapid, cost-effective model calls for analysis, coding, or multimodal understanding.
Qwen
Qwen3 VL 8B Instruct
Qwen3 VL 8B Instruct is an open-weight multimodal vision-language model developed by Qwen / Alibaba Cloud as part of the Qwen3-VL series, designed for instruction-following tasks that combine text with visual inputs such as images and video. Released around October 2025 under the Apache-2.0 license, it targets developers who need capable multimodal reasoning without the scale or cost of very large models.The model contains roughly 8.8 billion dense parameters and supports text, image, and video understanding with strong spatial perception, visual reasoning, and emerging visual agent abilities such as GUI interaction. A standout feature is its native ~256K token context window, extendable to around 1M tokens, enabling long-document reading and extended video comprehension. In today’s landscape, it balances openness, long-context capacity, and solid multimodal performance against heavier proprietary models. Typical applications include multimodal assistants, document and video analysis, visual question answering, and research or product prototyping where transparency and deployability matter.

Other OpenAI GPT Nano models

Other versions in the same family as GPT-4.1 nano.

GPT-4.1 nano License

Proprietary

License terms and commercial-use guidance for GPT-4.1 nano.

This model is proprietary. The author retains all rights, and use of the model is governed by their specific terms of service or license agreement.

Commercial use depends on the terms set by the model author. Most proprietary commercial models require a paid subscription, API key, or per-call billing. Check the provider’s pricing and terms-of-service for details.

License information is provided as a guide and is not legal advice.

Frequently Asked Questions About GPT-4.1 nano Vision

Yes. GPT-4.1 nano accepts image input, and on Roboflow's previous vision benchmark it passed 40.3% of visual understanding tasks (#69 of 77) and scored 57.2% on OCR.

GPT-4.1 nano has been deprecated by its provider and can no longer be run, so it is not part of Roboflow's current Vision Evals. Its results from the previous benchmark are preserved on this page for reference.