Roboflow

Captioning Model Rankings

Updated Aug 15

Find the best AI models for generating captions. Explore rankings to see which models describe images most accurately.

Vote in Arena
Votes power rankings.
Top Model Scores

ELO ratings for the highest performing models

Performance vs Accuracy

ELO score vs average latency • Better models are top-left

Action
1
multimodal1287520.99sGoogle
2
multimodal1246311.54sQwen
3
multimodal1245510.85sGoogle
4
multimodal1245312.32sQwen
5
multimodal1224310.42sGoogle
6
multimodal122338.91sGoogle
7
multimodal1223228.26sQwen
8
OpenAI
multimodal122359.94sOpenAI
9
multimodal1222322.27sQwen
10
multimodal121357.16sGoogle
11
multimodal121244.13sOpenAI
12
multimodal121255.13sAnthropic
13
OpenAI
multimodal121259.90sOpenAI
14
multimodal1212510.29sAnthropic
15
multimodal121238.79sAnthropic
16
multimodal1211517.18sGoogle
17
multimodal121136.57sQwen
18
multimodal120354.76sOpenAI
19
Grok
multimodal1203314.47sSpaceXAI
20
multimodal120025.35sOpenAI
21
multimodal1200315.05sMistral
22
multimodal120055.36sAnthropic
23
multimodal1200316.95sQwen
24
multimodal1200331.81sQwen
25
multimodal1200512.71sOpenAI
26
multimodal1199570.55sGoogle
27
multimodal119846.71sGoogle
28
multimodal1196310.69sGoogle
29
OpenAI
multimodal1192518.48sOpenAI
30
multimodal119056.20sAnthropic
31
multimodal119034.09sMistral
32
multimodal118933.35sQwen
33
multimodal1188417.97sQwen
34
multimodal1188323.40sQwen
35
OpenAI
multimodal118856.04sOpenAI
36
multimodal118836.08sQwen
37
multimodal1188414.42sQwen
38
OpenAI
multimodal118859.79sOpenAI
39
multimodal1188322.56sMeta
40
multimodal118734.80sMeta
41
multimodal118658.95sAnthropic
42
multimodal117952.10sGoogle
43
multimodal117636.00sMistral
44
multimodal116835.63sMeta
45
multimodal116756.65sAnthropic
46
multimodal115436.34sQwen
47
multimodal1147516.64sOpenAI
48
multimodal114435.94sMicrosoft

What is Captioning?

Image captioning generates a natural language description of an image. The model looks at a photo and produces a sentence or paragraph describing what it sees.

Rankings on this page are based on human preference votes in the Captioning Arena, where users pick which model's caption is more accurate and useful.

Frequently Asked Questions

Generating alt text for accessibility, tagging and searching media libraries, product description generation for e-commerce, content moderation pipelines, and as a preprocessing step feeding into larger AI workflows.

How is captioning different from open prompt?

Captioning generates a general description of an image with no input from you. Open prompt lets you ask specific questions or direct the model to focus on something particular.

Rankings are based on human preference votes in the Captioning Arena. Users see two captions side by side and vote on which is more accurate and useful.

Yes. Open the Captioning Playground, select the models you want to compare, and upload an image.