Compare the best 85 object detection models and try 60 of them on your own image, free in the Roboflow Playground. 49 are open-weight, so you can self-host them for free under their licenses.
85 models · 49 open-weight · 60 free to try · prices synced Jul 25, 2026
The best object detection model on our object detection benchmark right now is Gemini 3.5 Flash by Google, scoring 68.7% across 21 models tested, followed by GPT-5.6 Sol at 68.2%. Updated Jul 24, 2026. See the full ranking.
Benchmark Leader
Gemini 3.5 Flash
Highest score on our object detection benchmark, 21 models tested
Open-Weight Leader
Qwen3.5 27B
Best score among the open-weight models, #7 of 21 overall on our object detection benchmark
Fastest on Benchmark
Gemini 3.5 Flash-Lite
Quickest in our object detection benchmark runs: 4.1s per request on average
Lowest Measured Cost
Qwen3-VL 235B
Cheapest in our object detection benchmark runs: $0.0009 per request as measured, not list price
Ranked on our object detection benchmark, 21 models tested, updated Jul 24, 2026
49 models with downloadable weights you can self-host under their licenses (Modified MIT, Apache 2.0, and Custom). 24 run live in the Playground through hosted APIs, so self-hosting is optional.
Kimi K3Moonshot AI NEW | 2.8T | Modified MIT | Jul 2026 | |
Qwen3.6 35B A3BQwen | 35B total, 3B active | Apache 2.0 | Apr 2026 | |
Gemma 4 26B A4BGoogle | 25.2B | Apache 2.0 | Apr 2026 | |
Gemma 4 31BGoogle | 31B | Apache 2.0 | Apr 2026 | |
Qwen3.5 9bQwen | 9B | Apache 2.0 | Mar 2026 | |
| 122B | Apache 2.0 | Feb 2026 | ||
Qwen3.5 27BQwen | 27B | Apache 2.0 | Feb 2026 | |
Qwen3.5 35B A3BQwen | 35B | Apache 2.0 | Feb 2026 | |
| 397B | Apache 2.0 | Feb 2026 | ||
SAM 3Meta | — | Custom | Nov 2025 | |
| 8.8B | Apache 2.0 | Oct 2025 | ||
| 31B | Apache 2.0 | Oct 2025 | ||
YOLO26Ultralytics | 2.4M-55.7M | AGPL 3.0 | Oct 2025 | |
| 235B | Apache 2.0 | Sep 2025 | ||
Florence-2Microsoft | 230M | MIT | Jun 2025 | |
Llama 4 MaverickMeta | 400B | Proprietary | Apr 2025 | |
Llama 4 ScoutMeta | 109B | Proprietary | Apr 2025 | |
RF-DETRRoboflow | 30.5M-126.9M | Apache 2.0 | Mar 2025 | |
YOLOETHU-MIG | 10M-50M | AGPL 3.0 | Mar 2025 | |
YOLOv12THU-MIG | 2.6M-59.1M | AGPL 3.0 | Feb 2025 | |
| 7B | Apache 2.0 | Jan 2025 | ||
DEIMIntellindust AI Lab | 4M-62M | Apache 2.0 | Dec 2024 | |
D-FINEUSTC | 4M-62M | Apache 2.0 | Oct 2024 | |
YOLO11Ultralytics | 2.6M-56.9M | AGPL 3.0 | Sep 2024 | |
YOLOv10THU-MIG | 2.3M-29.5M | AGPL 3.0 | May 2024 | |
YOLO WorldTencent AI Lab | 13M | GPL v3 | Feb 2024 | |
YOLOv9Academia Sinica | 2.0M-57.3M | GPL v3 | Feb 2024 | |
Moondream 2Moondream | ~2B | Apache 2.0 | Jan 2024 | |
Grounded SAMIDEA Research | — | Apache 2.0 | Jan 2024 | |
YOLO-NASDeci AI | — | Custom | May 2023 | |
RT-DETRBaidu | 20M-76M | Apache 2.0 | Apr 2023 | |
Grounding DINOIDEA Research | 172M-341M | Apache 2.0 | Mar 2023 | |
YOLOv8Ultralytics | 3.2M-68.2M | AGPL 3.0 | Jan 2023 | |
RTMDetOpenMMLab | 4.8M-94.9M | GPL v3 | Dec 2022 | |
Co-DETROpenMMLab | 304M | MIT | Nov 2022 | |
YOLOv7Academia Sinica | 6.2M-151.7M | GPL v3 | Jul 2022 | |
OWL-ViTGoogle | — | Apache 2.0 | May 2022 | |
YOLOXMegvii | 0.91M-99.1M | Apache 2.0 | Jul 2021 | |
YOLOSHugging Face | — | MIT | Jun 2021 | |
YOLOv4-tinyAcademia Sinica | — | Custom | Nov 2020 | |
DETRMeta | ~41M | Apache 2.0 | May 2020 | |
YOLOv4Academia Sinica | — | — | Apr 2020 | |
YOLOv5Ultralytics | 1.9M-86.7M | AGPL 3.0 | Jan 2020 | |
EfficientDetGoogle | 3.9M-51.9M | Apache 2.0 | Nov 2019 | |
Detectron2Meta | — | Apache 2.0 | Sep 2019 | |
MediaPipeGoogle | — | Apache 2.0 | Jul 2019 | |
MobileNet SSD v2Google | 15.3M | MIT | Jan 2018 | |
Mask R-CNNMeta | 44.4M | MIT | Oct 2017 | |
Faster R-CNNMicrosoft | 41.8M | MIT | Jun 2015 |
36 proprietary models where the weights aren't downloadable: access is through each provider's API and billed by them. Try all of them free in the Playground.
Claude Opus 5Anthropic NEW | $5.00 | $25.00 | 1M | Jul 2026 | |
Gemini 3.5 Flash-LiteGoogle NEW | $0.30 | $2.50 | 1.0M | Jul 2026 | |
Gemini 3.6 FlashGoogle NEW | $1.50 | $7.50 | 1M | Jul 2026 | |
GPT-5.6 LunaOpenAI NEW | $1.00 | $6.00 | 1.5M | Jul 2026 | |
GPT-5.6 SolOpenAI NEW | $5.00 | $30.00 | 1.5M | Jul 2026 | |
GPT-5.6 TerraOpenAI NEW | $2.50 | $15.00 | 1.1M | Jul 2026 | |
Muse Spark 1.1Meta NEW | $1.25 | $4.25 | 1.0M | Jul 2026 | |
Claude Sonnet 5Anthropic NEW | $2.00 | $10.00 | 1M | Jun 2026 | |
Claude Fable 5Anthropic NEW | $10.00 | $50.00 | 1M | Jun 2026 | |
Claude Opus 4.8Anthropic | $5.00 | $25.00 | 1M | May 2026 | |
Gemini 3.5 FlashGoogle | $1.50 | $9.00 | 1.0M | May 2026 | |
GPT-5.5OpenAI | $5.00 | $30.00 | 1M | Apr 2026 | |
Claude Opus 4.7Anthropic | $5.00 | $25.00 | 1M | Apr 2026 | |
Qwen3.6 PlusQwen | $0.33 | $1.95 | 1M | Apr 2026 | |
GPT-5.4 MiniOpenAI | $0.75 | $4.50 | 400K | Mar 2026 | |
GPT-5.4 NanoOpenAI | $0.20 | $1.25 | 400K | Mar 2026 | |
GPT-5.4OpenAI | $2.50 | $15.00 | 1.1M | Mar 2026 | |
Gemini 3.1 Flash-LiteGoogle | $0.25 | $1.50 | 1M | Mar 2026 | |
Gemini 3.1 ProGoogle | $2.00 | $12.00 | 1M | Feb 2026 | |
Claude Sonnet 4.6Anthropic | $3.00 | $15.00 | 1M | Feb 2026 | |
Claude Opus 4.6 Anthropic | $5.00 | $25.00 | 1M | Feb 2026 | |
Gemini 3 FlashGoogle | $0.50 | $3.00 | 1M | Dec 2025 | |
GPT-5.2OpenAI | $1.75 | $14.00 | 400K | Dec 2025 | |
Claude Opus 4.5Anthropic | $5.00 | $25.00 | 200K | Nov 2025 | |
GPT-5.1OpenAI | $1.25 | $10.00 | 196K | Nov 2025 | |
Claude Haiku 4.5Anthropic | $1.00 | $5.00 | 200K | Oct 2025 | |
Claude Sonnet 4.5Anthropic | $3.00 | $15.00 | 200K | Sep 2025 | |
GPT-5OpenAI | $1.25 | $10.00 | — | Aug 2025 | |
GPT-5 MiniOpenAI | $0.25 | $2.00 | 400K | Aug 2025 | |
GPT-5 NanoOpenAI | $0.050 | $0.40 | 400K | Aug 2025 | |
Claude Opus 4.1Anthropic | $15.00 | $75.00 | 200K | Aug 2025 | |
Gemini 2.5 Flash-LiteGoogle | $0.10 | $0.40 | 1M | Jul 2025 | |
Gemini 2.5 FlashGoogle | $0.30 | $2.50 | 1M | Jul 2025 | |
Grok 4xAI | — | — | — | Jul 2025 | |
Gemini 2.5 ProGoogle | $1.25 | $10.00 | 1M | Jun 2025 | |
Qwen VL MaxQwen | — | — | — | Feb 2025 |
The right object detection model depends less on which one tops a benchmark and more on three things about your problem: whether your classes are fixed or open-ended, how much labeled data you have, and where the model has to run. The models on this page fall into three families that answer those questions differently.
Architectures like RF-DETR, the YOLO line, and RT-DETR are built to do one thing well: detect a fixed set of classes fast and accurately. Fine-tuned on a few hundred to a few thousand labeled images of your own objects, a specialized detector is almost always the most accurate, cheapest, and lowest-latency option for a production task with stable classes. It runs in milliseconds on a GPU and can be deployed to the edge.
The cost is up front: you need labeled data and a training step. Choose this family when the classes are known and fixed (a fixed set of products, defects, or vehicle types), accuracy on your specific domain matters, or you are detecting thousands of images and per-call API cost or latency would add up.
Models like Grounding DINO and YOLO-World detect objects from a text prompt at inference time, with no training. Ask for "forklift" or "person wearing a hard hat" and they will try to find it. This is the right family when your categories change often, you have no labeled data yet, or you are validating an idea before investing in a dataset. The tradeoff is accuracy on fine-grained or unusual classes and higher latency than a fine-tuned detector.
General-purpose VLMs (GPT, Claude, Gemini, Qwen VL) can localize objects from a prompt and, unlike a detector, also reason about the scene, read text, and answer questions about what they find. They are the most flexible option and need no training, but they are the slowest and most expensive per image and are generally less precise at tight bounding boxes than a purpose-built detector. Reach for a VLM when detection is one step in a broader understanding task, or when you need a description alongside the boxes.
These families are not mutually exclusive. A frequent setup prototypes with a VLM or a zero-shot detector to validate the task and generate initial labels, then trains a specialized detector once the classes settle and volume grows. In a Roboflow Workflow you can also chain them: a fast detector localizes the objects, and a VLM interprets only the crops that matter, which keeps cost down while adding reasoning where it counts.
The bottom line: Fixed classes and production volume point to a fine-tuned specialized detector; changing or unknown classes point to a zero-shot detector or a VLM. Prototype with the flexible option, then move to a trained model when accuracy and cost start to matter.
Object detection is the computer vision task of finding every object of interest in an image and returning a class label, a confidence score, and a bounding box for each one. Unlike classification, which produces one label for the whole image, detection answers what is present, how many, and where. Modern detectors fall into three families: single-stage models like the YOLO line that predict boxes in one pass for speed, two-stage models like Faster R-CNN that refine proposed regions for accuracy, and transformer set-prediction models like DETR and RF-DETR that learn to output a clean set of boxes without hand-tuned anchors or NMS. Vision language models can also detect zero-shot from a text prompt, trading precision for flexibility. Accuracy is measured by mean average precision (mAP) at IoU thresholds, and the task underpins counting, safety monitoring, shelf audits, and autonomous systems. This page lists 85 object detection models, including 49 open-weight options you can self-host; 60 of them run live in the Playground so you can test them on your own images.
On our object detection benchmark (21 models tested, updated Jul 24, 2026), Gemini 3.5 Flash by Google currently scores highest at 68.7%. The full ranking is on our evals page. For fixed categories in production, a model fine-tuned on your own data still often wins.
Yes. 49 of the 85 object detection models here are open-weight (for example Kimi K3, Qwen3.6 35B A3B, and Gemma 4 26B A4B), free to self-host under their licenses (Modified MIT, Apache 2.0, and Custom).
Yes. You can run 60 of them in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
This page lists all 85 object detection models in the Roboflow Playground catalog: 49 open-weight models you can self-host and 36 proprietary models accessed through provider APIs; 60 of them run live in the Roboflow Playground on your own images. On our object detection benchmark (21 models tested), Gemini 3.5 Flash by Google currently scores highest. Compare licenses, parameters, API prices, and release dates side by side, or open any model page for full details.