Captioning Model Rankings
Updated Aug 15Find the best AI models for generating captions. Explore rankings to see which models describe images most accurately.
ELO ratings for the highest performing models
ELO score vs average latency • Better models are top-left
| Action | |||||||
|---|---|---|---|---|---|---|---|
1 | multimodal | 1287 | 5 | 20.99s | |||
2 | multimodal | 1246 | 3 | 11.54s | Qwen | ||
3 | multimodal | 1245 | 5 | 10.85s | |||
4 | multimodal | 1245 | 3 | 12.32s | Qwen | ||
5 | multimodal | 1224 | 3 | 10.42s | |||
6 | multimodal | 1223 | 3 | 8.91s | |||
7 | multimodal | 1223 | 2 | 28.26s | Qwen | ||
8 | multimodal | 1223 | 5 | 9.94s | OpenAI | ||
9 | multimodal | 1222 | 3 | 22.27s | Qwen | ||
10 | multimodal | 1213 | 5 | 7.16s | |||
11 | multimodal | 1212 | 4 | 4.13s | OpenAI | ||
12 | multimodal | 1212 | 5 | 5.13s | Anthropic | ||
13 | multimodal | 1212 | 5 | 9.90s | OpenAI | ||
14 | multimodal | 1212 | 5 | 10.29s | Anthropic | ||
15 | multimodal | 1212 | 3 | 8.79s | Anthropic | ||
16 | multimodal | 1211 | 5 | 17.18s | |||
17 | multimodal | 1211 | 3 | 6.57s | Qwen | ||
18 | multimodal | 1203 | 5 | 4.76s | OpenAI | ||
19 | multimodal | 1203 | 3 | 14.47s | SpaceXAI | ||
20 | multimodal | 1200 | 2 | 5.35s | OpenAI | ||
21 | multimodal | 1200 | 3 | 15.05s | Mistral | ||
22 | multimodal | 1200 | 5 | 5.36s | Anthropic | ||
23 | multimodal | 1200 | 3 | 16.95s | Qwen | ||
24 | multimodal | 1200 | 3 | 31.81s | Qwen | ||
25 | multimodal | 1200 | 5 | 12.71s | OpenAI | ||
26 | multimodal | 1199 | 5 | 70.55s | |||
27 | multimodal | 1198 | 4 | 6.71s | |||
28 | multimodal | 1196 | 3 | 10.69s | |||
29 | multimodal | 1192 | 5 | 18.48s | OpenAI | ||
30 | multimodal | 1190 | 5 | 6.20s | Anthropic | ||
31 | multimodal | 1190 | 3 | 4.09s | Mistral | ||
32 | multimodal | 1189 | 3 | 3.35s | Qwen | ||
33 | multimodal | 1188 | 4 | 17.97s | Qwen | ||
34 | multimodal | 1188 | 3 | 23.40s | Qwen | ||
35 | multimodal | 1188 | 5 | 6.04s | OpenAI | ||
36 | multimodal | 1188 | 3 | 6.08s | Qwen | ||
37 | multimodal | 1188 | 4 | 14.42s | Qwen | ||
38 | multimodal | 1188 | 5 | 9.79s | OpenAI | ||
39 | multimodal | 1188 | 3 | 22.56s | Meta | ||
40 | multimodal | 1187 | 3 | 4.80s | Meta | ||
41 | multimodal | 1186 | 5 | 8.95s | Anthropic | ||
42 | multimodal | 1179 | 5 | 2.10s | |||
43 | multimodal | 1176 | 3 | 6.00s | Mistral | ||
44 | multimodal | 1168 | 3 | 5.63s | Meta | ||
45 | multimodal | 1167 | 5 | 6.65s | Anthropic | ||
46 | multimodal | 1154 | 3 | 6.34s | Qwen | ||
47 | multimodal | 1147 | 5 | 16.64s | OpenAI | ||
48 | multimodal | 1144 | 3 | 5.94s | Microsoft |
What is Captioning?
Image captioning generates a natural language description of an image. The model looks at a photo and produces a sentence or paragraph describing what it sees.
Rankings on this page are based on human preference votes in the Captioning Arena, where users pick which model's caption is more accurate and useful.
Frequently Asked Questions
Generating alt text for accessibility, tagging and searching media libraries, product description generation for e-commerce, content moderation pipelines, and as a preprocessing step feeding into larger AI workflows.
How is captioning different from open prompt?
Captioning generates a general description of an image with no input from you. Open prompt lets you ask specific questions or direct the model to focus on something particular.
Rankings are based on human preference votes in the Captioning Arena. Users see two captions side by side and vote on which is more accurate and useful.
Yes. Open the Captioning Playground, select the models you want to compare, and upload an image.