Compare the best 82 image captioning models and try 76 of them on your own images, free in the Roboflow Playground. 32 are open-weight, so you can self-host them for free under their licenses.
82 models · 32 open-weight · 76 free to try · prices synced Sep 12, 2026
We haven't benchmarked image captioning yet; scores and rankings appear on this site only where we've measured them. Until then, compare the models below, and for fixed categories in production expect a model fine-tuned on your own data to win.
32 models with downloadable weights you can self-host under their licenses (MIT, Apache 2.0, and Modified MIT). 26 run live in the Playground through hosted APIs, so self-hosting is optional.
| Actions | ||||
|---|---|---|---|---|
GLM 5.3 FlashZ.ai | 320B total, 18B active | MIT | Aug 2026 | |
Qwen3.8 27BQwen | 27.78B | Apache 2.0 | Aug 2026 | |
Muse Glimmer 30BMeta | 29.6B | Apache 2.0 | Aug 2026 | |
Kimi K3Moonshot AI | 2.8T | Modified MIT | Jul 2026 | |
Gemma 4 12BGoogle | 12B | Apache 2.0 | Jun 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
Qwen3.5-27BQwen | 27B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Custom | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Custom | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Custom | Mar 2025 | |
Gemma 3 27BGoogle | — | Custom | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Custom | Mar 2025 | |
SmolVLM2Hugging Face | 256M – 2.2B | Apache 2.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
PaliGemma 2Google | 3B, 10B, 28B | Custom | Dec 2024 | |
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 | |
PaliGemmaGoogle | 3B | Custom | May 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
Qwen-VLQwen | — | Custom | Aug 2023 |
50 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
| Actions | |||||
|---|---|---|---|---|---|
GPT-6 AstraOpenAI NEW | $10.00 | $50.00 | 1.1M | Sep 2026 | |
Gemini 3.8 FlashGoogle NEW | $0.75 | $3.75 | 1.0M | Sep 2026 | |
Muse Spark 1.3Meta NEW | $1.25 | $4.25 | 1.0M | Sep 2026 | |
Claude Fable 5.1Anthropic NEW | $10.00 | $50.00 | 1M | Sep 2026 | |
Qwen3.8 FlashQwen | $0.15 | $0.47 | 1M | Aug 2026 | |
Gemini 3.7 FlashGoogle | $0.75 | $3.75 | 1.0M | Aug 2026 | |
Grok 4.6SpaceXAI | $2.00 | $6.00 | 500K | Aug 2026 | |
Muse Spark 1.2Meta | $1.25 | $4.25 | 1.0M | Aug 2026 | |
Qwen3.8 MaxQwen | — | — | 984K | Aug 2026 | |
Qwen3.7 FlashQwen | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle | $0.75 | $3.75 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI | $0.20 | $1.20 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI | $2.00 | $10.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI | $2.00 | $12.00 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Grok 4.5SpaceXAI | $2.00 | $6.00 | 500K | Jul 2026 | |
Claude Sonnet 5Anthropic | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic | $10.00 | $50.00 | 1M | Jun 2026 | |
Qwen3.7 PlusQwen | $0.32 | $1.28 | — | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GLM 5V TurboZ.ai | $1.20 | $4.00 | 200K | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4SpaceXAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
Captioning is an encoder-decoder problem that modern VLMs handle natively, so the real choice is not which architecture but which tier: how much caption quality you need, at what volume, and whether captions are the end product or an input to another system.
The largest models write the most detailed, specific, and correct captions, follow style instructions (length, tone, what to emphasize), and can caption dense or unusual scenes without falling apart. They cost the most per image and run the slowest, which matters when you caption millions of assets rather than dozens.
Compact open models (Gemma-class VLMs, Florence-2 with its dense region captioning) produce serviceable captions at a fraction of the cost, can be self-hosted for privacy or offline use, and are often the right economics for alt-text generation and bulk media indexing where a plain, accurate sentence is enough.
Short captions suit alt text and search indexing; detailed captions suit retrieval systems and datasets for training other models, but longer generations also hallucinate more specifics, so validate on your own images. If captions feed a search index, consistency across the library usually matters more than peak eloquence, which favors the cheaper tier.
The bottom line: Caption quality as the product: use a frontier VLM. Captions as infrastructure at volume: use a small open model and spend the savings on coverage.
Image captioning is the task of generating a natural language description of an image. Architecturally it is an encoder-decoder problem: a vision encoder turns the image into features, and a language decoder writes text conditioned on them, which is exactly the structure of modern vision language models. Variants include short captions (one summary sentence), detailed captions, and dense captioning that describes individual regions. Automatic metrics like CIDEr and BLEU exist but mostly reward phrasing overlap with reference captions, so human or model-based judgment is common in practice. Captioning is used for alt text and accessibility, media library indexing, and making image collections searchable by content. This page lists 82 image captioning models, including 32 open-weight options you can self-host; 76 of them run live in the Playground so you can test them on your own images.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the image captioning models on this page and try them on your own images to see which fits.
Yes. 32 of the 82 image captioning models here are open-weight (for example GLM 5.3 Flash, Qwen3.8 27B, and Muse Glimmer 30B), free to self-host under their licenses (MIT, Apache 2.0, and Modified MIT).
Yes. You can run 76 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 82 image captioning models in the Roboflow Playground catalog: 32 open-weight models you can self-host and 50 proprietary models accessed through provider APIs; 76 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.