Compare the best 74 multi-label classification models and try them on your own images, free in the Roboflow Playground. 25 are open-weight, so you can self-host them for free under their licenses.
74 models · 25 open-weight · 74 free to try · prices synced Sep 12, 2026
We haven't benchmarked multi-label classification yet; scores and rankings appear on this site only where we've measured them. Until then, compare the models below, and for fixed categories in production expect a model fine-tuned on your own data to win.
25 models with downloadable weights you can self-host under their licenses (MIT, Apache 2.0, and Modified MIT). All run live in the Playground through hosted APIs, so self-hosting is optional.
| Actions | ||||
|---|---|---|---|---|
GLM 5.3 FlashZ.ai | 320B total, 18B active | MIT | Aug 2026 | |
Qwen3.8 27BQwen | 27.78B | Apache 2.0 | Aug 2026 | |
Muse Glimmer 30BMeta | 29.6B | Apache 2.0 | Aug 2026 | |
Kimi K3Moonshot AI | 2.8T | Modified MIT | Jul 2026 | |
Qwen3.6 27BQwen | 27B | Apache 2.0 | Apr 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
Qwen3.5-27BQwen | 27B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
Kimi K2.5Moonshot AI | 1T | Modified MIT | Jan 2026 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
| 235B | Apache 2.0 | Sep 2025 | ||
Llama 4 MaverickMeta | 400B | Custom | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Custom | Apr 2025 | |
Mistral Small 3.1 24BMistral | 24B | Apache 2.0 | Mar 2025 | |
Gemma 3 12BGoogle | 12B | Custom | Mar 2025 | |
Gemma 3 27BGoogle | — | Custom | Mar 2025 | |
Gemma 3 4BGoogle | 4B | Custom | Mar 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
Pixtral 12BMistral | 12B | Apache 2.0 | Sep 2024 |
49 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
| Actions | |||||
|---|---|---|---|---|---|
GPT-6 AstraOpenAI NEW | $10.00 | $50.00 | 1.1M | Sep 2026 | |
Gemini 3.8 FlashGoogle NEW | $0.75 | $3.75 | 1.0M | Sep 2026 | |
Muse Spark 1.3Meta NEW | $1.25 | $4.25 | 1.0M | Sep 2026 | |
Claude Fable 5.1Anthropic NEW | $10.00 | $50.00 | 1M | Sep 2026 | |
Qwen3.8 FlashQwen | $0.15 | $0.47 | 1M | Aug 2026 | |
Gemini 3.7 FlashGoogle | $0.75 | $3.75 | 1.0M | Aug 2026 | |
Grok 4.6SpaceXAI | $2.00 | $6.00 | 500K | Aug 2026 | |
Muse Spark 1.2Meta | $1.25 | $4.25 | 1.0M | Aug 2026 | |
Qwen3.8 MaxQwen | — | — | 984K | Aug 2026 | |
Qwen3.7 FlashQwen | $0.030 | $0.13 | 1M | Jul 2026 | |
Claude Opus 5Anthropic | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle | $0.75 | $3.75 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI | $0.20 | $1.20 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI | $2.00 | $10.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI | $2.00 | $12.00 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Grok 4.5SpaceXAI | $2.00 | $6.00 | 500K | Jul 2026 | |
Claude Sonnet 5Anthropic | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 FlashQwen | $0.19 | $1.13 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GLM 5V TurboZ.ai | $1.20 | $4.00 | 200K | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
Mistral Medium 3.1Mistral | $0.40 | $2.00 | 128K | Aug 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4SpaceXAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
Real images rarely contain exactly one thing, and multi-label models score every label independently. The choice mirrors single-label classification, with one extra ingredient: thresholds.
CLIP-style scoring and VLM prompting assign multiple labels with no training and adapt instantly as the label set changes. They are the right start when the taxonomy is still moving or labeled data does not exist yet.
With a stable label set, a trained model with a sigmoid per label is cheaper, more consistent, and properly evaluable per label. The underrated work is threshold calibration: each label needs its own operating point tuned to your precision and recall targets, and co-occurrence patterns (street scenes imply vehicles) are something trained models learn and zero-shot ones miss.
The bottom line: Moving taxonomy: zero-shot or VLM labeling. Stable taxonomy at volume: train a multi-label classifier and spend real time calibrating per-label thresholds.
Multi-label classification is the task of assigning several labels to one image simultaneously, because real scenes rarely contain exactly one thing. Instead of a softmax that forces classes to compete, the model scores each label independently with a sigmoid, and per-label thresholds are tuned to balance precision and recall; label co-occurrence (street scenes imply vehicles) is part of what models learn. Performance is reported as per-label and averaged F1 or mAP. It is used for attribute tagging in e-commerce, radiology findings where multiple conditions coexist, and content screening across simultaneous policies. This page lists 74 multi-label classification models, including 25 open-weight options you can self-host; all of them run live in the Playground so you can test them on your own images.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the multi-label classification models on this page and try them on your own images to see which fits.
Yes. 25 of the 74 multi-label classification models here are open-weight (for example GLM 5.3 Flash, Qwen3.8 27B, and Muse Glimmer 30B), free to self-host under their licenses (MIT, Apache 2.0, and Modified MIT).
Yes. You can run all 74 in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 74 multi-label classification models in the Roboflow Playground catalog: 25 open-weight models you can self-host and 49 proprietary models accessed through provider APIs; all of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.