Open Prompt Model Rankings
Updated Aug 26See which models excel at free-form prompts. Compare performance, discover strengths, and find the right model for open-ended vision tasks.
ELO ratings for the highest performing models
ELO score vs average latency • Better models are top-left
| Action | |||||||
|---|---|---|---|---|---|---|---|
1 | multimodal | 1311 | 5 | 17.39s | |||
2 | multimodal | 1287 | 5 | 31.91s | OpenAI | ||
3 | multimodal | 1262 | 5 | 9.99s | |||
4 | multimodal | 1244 | 5 | 5.83s | OpenAI | ||
5 | multimodal | 1235 | 3 | 11.91s | Qwen | ||
6 | multimodal | 1225 | 3 | 34.69s | SpaceXAI | ||
7 | multimodal | 1224 | 2 | 4.28s | |||
8 | multimodal | 1223 | 3 | 40.38s | |||
9 | multimodal | 1222 | 3 | 14.38s | |||
10 | multimodal | 1222 | 4 | 8.99s | Qwen | ||
11 | multimodal | 1215 | 5 | 5.56s | Anthropic | ||
12 | multimodal | 1215 | 3 | 22.23s | Qwen | ||
13 | multimodal | 1211 | 5 | 27.67s | |||
14 | multimodal | 1211 | 4 | 2.46s | OpenAI | ||
15 | multimodal | 1211 | 5 | 8.00s | Anthropic | ||
16 | multimodal | 1210 | 3 | 7.10s | Qwen | ||
17 | multimodal | 1201 | 4 | 53.89s | Qwen | ||
18 | multimodal | 1200 | 2 | 22.55s | Moonshot AI | ||
19 | multimodal | 1200 | 5 | 15.36s | OpenAI | ||
20 | multimodal | 1200 | 4 | 3.42s | |||
21 | multimodal | 1200 | 2 | 37.07s | Meta | ||
22 | multimodal | 1200 | 5 | 88.03s | OpenAI | ||
23 | multimodal | 1200 | 2 | 11.52s | Anthropic | ||
24 | multimodal | 1200 | 4 | 6.25s | OpenAI | ||
25 | multimodal | 1200 | 5 | 17.02s | Anthropic | ||
26 | multimodal | 1200 | 3 | 15.41s | Qwen | ||
27 | multimodal | 1200 | 4 | 7.39s | Qwen | ||
28 | multimodal | 1200 | 2 | 27.80s | Qwen | ||
29 | multimodal | 1200 | 2 | 13.13s | Qwen | ||
30 | multimodal | 1200 | 3 | 9.25s | Qwen | ||
31 | multimodal | 1200 | 4 | 8.02s | Anthropic | ||
32 | multimodal | 1200 | 3 | 18.06s | Anthropic | ||
33 | multimodal | 1199 | 5 | 1.94s | |||
34 | multimodal | 1198 | 3 | 3.45s | Meta | ||
35 | multimodal | 1197 | 5 | 4.39s | OpenAI | ||
36 | multimodal | 1197 | 3 | 12.06s | Qwen | ||
37 | multimodal | 1193 | 5 | 20.49s | OpenAI | ||
38 | multimodal | 1192 | 3 | 13.57s | Qwen | ||
39 | multimodal | 1190 | 3 | 3.56s | Meta | ||
40 | multimodal | 1188 | 5 | 10.08s | |||
41 | multimodal | 1188 | 5 | 2.86s | OpenAI | ||
42 | multimodal | 1188 | 1 | 32.69s | Qwen | ||
43 | multimodal | 1188 | 5 | 1.92s | OpenAI | ||
44 | multimodal | 1188 | 5 | 11.29s | Anthropic | ||
45 | multimodal | 1188 | 2 | 29.98s | Qwen | ||
46 | multimodal | 1188 | 4 | 3.36s | Meta | ||
47 | multimodal | 1187 | 5 | 7.76s | OpenAI | ||
48 | multimodal | 1181 | 3 | 26.20s | |||
49 | multimodal | 1181 | 5 | 5.43s | |||
50 | multimodal | 1180 | 3 | 6.40s | Mistral | ||
51 | multimodal | 1177 | 5 | 7.80s | Anthropic | ||
52 | multimodal | 1167 | 3 | 5.72s | Qwen | ||
53 | multimodal | 1167 | 3 | 12.91s | Qwen | ||
54 | multimodal | 1157 | 3 | 9.70s | |||
55 | multimodal | 1148 | 3 | 7.41s | Mistral | ||
56 | multimodal | 1144 | 5 | 4.71s | Anthropic | ||
57 | multimodal | 1142 | 3 | 5.33s | Mistral |
What is Open Prompt?
Open prompt lets you ask a model anything about an image in plain language. Instead of running a fixed task, you write a custom prompt and the model interprets the image based on your instructions.
This makes it useful for tasks that don't fit a predefined category: pulling structured data from a photo, counting specific items, describing spatial relationships, reading charts, or flagging anomalies.
Rankings on this page are based on ELO scores from the Open Prompt Arena, run across a diverse mix of user-submitted tasks. A model that excels at one type of prompt may underperform on another. The score reflects overall breadth.
Frequently Asked Questions
Extracting structured data from photos, counting specific objects in a scene, reading and interpreting charts or graphs, identifying anomalies in images, describing spatial relationships, and answering questions about image content in natural language.
Insurance for claims assessment from photos, retail for shelf and planogram analysis, construction and real estate for site documentation, research for data extraction from visual materials, and customer support for image-based troubleshooting.
OCR is purpose-built for extracting text from images and is generally more accurate on that specific task. Open prompt is more flexible and can handle a wider range of questions, but for pure text extraction OCR models will perform better.
Rankings are based on ELO scores from head-to-head battles in the Open Prompt Arena. Users submit their own prompts and images, then vote on which model's response is more accurate or useful.
Yes. Open the Open Prompt Playground, select the models you want to compare, upload an image, and type your question or instruction.