Roboflow

Open Prompt Model Rankings

Updated Aug 26

See which models excel at free-form prompts. Compare performance, discover strengths, and find the right model for open-ended vision tasks.

Vote in Arena
Votes power rankings.
Top Model Scores

ELO ratings for the highest performing models

Performance vs Accuracy

ELO score vs average latency • Better models are top-left

Action
1
multimodal1311517.39sGoogle
2
multimodal1287531.91sOpenAI
3
multimodal126259.99sGoogle
4
OpenAI
multimodal124455.83sOpenAI
5
multimodal1235311.91sQwen
6
Grok
multimodal1225334.69sSpaceXAI
7
multimodal122424.28sGoogle
8
multimodal1223340.38sGoogle
9
multimodal1222314.38sGoogle
10
multimodal122248.99sQwen
11
multimodal121555.56sAnthropic
12
multimodal1215322.23sQwen
13
multimodal1211527.67sGoogle
14
multimodal121142.46sOpenAI
15
multimodal121158.00sAnthropic
16
multimodal121037.10sQwen
17
multimodal1201453.89sQwen
18
MoonshotAI
multimodal1200222.55sMoonshot AI
19
OpenAI
multimodal1200515.36sOpenAI
20
multimodal120043.42sGoogle
21
multimodal1200237.07sMeta
22
OpenAI
multimodal1200588.03sOpenAI
23
multimodal1200211.52sAnthropic
24
multimodal120046.25sOpenAI
25
multimodal1200517.02sAnthropic
26
multimodal1200315.41sQwen
27
multimodal120047.39sQwen
28
multimodal1200227.80sQwen
29
multimodal1200213.13sQwen
30
multimodal120039.25sQwen
31
multimodal120048.02sAnthropic
32
Anthropic
multimodal1200318.06sAnthropic
33
multimodal119951.94sGoogle
34
multimodal119833.45sMeta
35
OpenAI
multimodal119754.39sOpenAI
36
multimodal1197312.06sQwen
37
multimodal1193520.49sOpenAI
38
multimodal1192313.57sQwen
39
multimodal119033.56sMeta
40
multimodal1188510.08sGoogle
41
multimodal118852.86sOpenAI
42
multimodal1188132.69sQwen
43
multimodal118851.92sOpenAI
44
multimodal1188511.29sAnthropic
45
multimodal1188229.98sQwen
46
multimodal118843.36sMeta
47
OpenAI
multimodal118757.76sOpenAI
48
multimodal1181326.20sGoogle
49
multimodal118155.43sGoogle
50
multimodal118036.40sMistral
51
multimodal117757.80sAnthropic
52
multimodal116735.72sQwen
53
multimodal1167312.91sQwen
54
multimodal115739.70sGoogle
55
multimodal114837.41sMistral
56
multimodal114454.71sAnthropic
57
multimodal114235.33sMistral

What is Open Prompt?

Open prompt lets you ask a model anything about an image in plain language. Instead of running a fixed task, you write a custom prompt and the model interprets the image based on your instructions.

This makes it useful for tasks that don't fit a predefined category: pulling structured data from a photo, counting specific items, describing spatial relationships, reading charts, or flagging anomalies.

Rankings on this page are based on ELO scores from the Open Prompt Arena, run across a diverse mix of user-submitted tasks. A model that excels at one type of prompt may underperform on another. The score reflects overall breadth.

Frequently Asked Questions

Extracting structured data from photos, counting specific objects in a scene, reading and interpreting charts or graphs, identifying anomalies in images, describing spatial relationships, and answering questions about image content in natural language.

Insurance for claims assessment from photos, retail for shelf and planogram analysis, construction and real estate for site documentation, research for data extraction from visual materials, and customer support for image-based troubleshooting.

OCR is purpose-built for extracting text from images and is generally more accurate on that specific task. Open prompt is more flexible and can handle a wider range of questions, but for pure text extraction OCR models will perform better.

Rankings are based on ELO scores from head-to-head battles in the Open Prompt Arena. Users submit their own prompts and images, then vote on which model's response is more accurate or useful.

Yes. Open the Open Prompt Playground, select the models you want to compare, upload an image, and type your question or instruction.