Compare the best 68 image captioning models and try 62 of them on your own image, free in the Roboflow Playground. 29 are open-weight, so you can self-host them for free under their licenses.
68 models · 29 open-weight · 62 free to try · prices synced Jul 28, 2026
We haven't benchmarked image captioning yet; scores and rankings appear on this site only where we've measured them. Until then, compare the models below, and for fixed categories in production expect a model fine-tuned on your own data to win.
29 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and MIT). 23 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle NEW | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Proprietary | Mar 2025 | |
Gemma 3 27BGoogle | — | Proprietary | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Proprietary | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
Qwen-VLQwen | — | Custom | Aug 2023 |
39 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Qwen3.7 FlashQwen NEW | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $0.50 | $3.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $1.25 | $7.50 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
Captioning is an encoder-decoder problem that modern VLMs handle natively, so the real choice is not which architecture but which tier: how much caption quality you need, at what volume, and whether captions are the end product or an input to another system.
The largest models write the most detailed, specific, and correct captions, follow style instructions (length, tone, what to emphasize), and can caption dense or unusual scenes without falling apart. They cost the most per image and run the slowest, which matters when you caption millions of assets rather than dozens.
Compact open models (Gemma-class VLMs, Florence-2 with its dense region captioning) produce serviceable captions at a fraction of the cost, can be self-hosted for privacy or offline use, and are often the right economics for alt-text generation and bulk media indexing where a plain, accurate sentence is enough.
Short captions suit alt text and search indexing; detailed captions suit retrieval systems and datasets for training other models, but longer generations also hallucinate more specifics, so validate on your own images. If captions feed a search index, consistency across the library usually matters more than peak eloquence, which favors the cheaper tier.
The bottom line: Caption quality as the product: use a frontier VLM. Captions as infrastructure at volume: use a small open model and spend the savings on coverage.
Image captioning is the task of generating a natural language description of an image. Architecturally it is an encoder-decoder problem: a vision encoder turns the image into features, and a language decoder writes text conditioned on them, which is exactly the structure of modern vision language models. Variants include short captions (one summary sentence), detailed captions, and dense captioning that describes individual regions. Automatic metrics like CIDEr and BLEU exist but mostly reward phrasing overlap with reference captions, so human or model-based judgment is common in practice. Captioning is used for alt text and accessibility, media library indexing, and making image collections searchable by content. This page lists 68 image captioning models, including 29 open-weight options you can self-host; 62 of them run live in the Playground so you can test them on your own images.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the image captioning models on this page and try them on your own image to see which fits.
Yes. 29 of the 68 image captioning models here are open-weight (for example Kimi K3, Gemma 4 12B, and Qwen3.6 27B), free to self-host under their licenses (Modified MIT, Apache 2.0, and MIT).
Yes. You can run 62 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 68 image captioning models in the Roboflow Playground catalog: 29 open-weight models you can self-host and 39 proprietary models accessed through provider APIs; 62 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.