Compare all 5 Microsoft vision models we track, all of them open-weight. Try 1 of them on your own images, free in the Roboflow Playground.
5 models · 5 open-weight · 1 free to try · Updated Jun 2025
Every Microsoft vision model in our catalog is open-weight. All 5 publish downloadable weights, so you can self-host under their own licenses instead of paying per request. The open-weight side runs on MIT and Custom licenses. The newest additions are Florence-2 (Jun 2025) and LLaVA-1.5 (Oct 2023).
Between them, the Microsoft models we list cover 10 distinct vision tasks. The widest coverage is object detection (2 models), OCR (2), and captioning (1). For object detection, start with Florence-2 or Faster R-CNN.
On the open-weight tier, LLaVA-1.5 is the largest at 7B, 13B parameters and ResNet-50 the smallest at 25.6M, small enough to run on a single GPU or on-device. 4 of the 5 open Microsoft models carry a permissive license, so commercial use is straightforward; the other 1 ship under terms worth reading before you deploy.
With every model here open-weight, the real trade is size against generality: the smaller checkpoints fine-tune and deploy cheaply on your own hardware, while the larger ones cover more ground out of the box. 1 of the 5 Microsoft vision models run live in the Roboflow Playground, so you can run the same image through several of them and compare the answers before committing to one.
What each of the 5 Microsoft vision models in our catalog is built for, and how you run it.
5 models with downloadable weights you can self-host under their licenses (MIT and Custom). 1 run live in the Playground through hosted APIs, so self-hosting is optional.
2 of the 5 Microsoft vision models we track handle object detection: Florence-2 and Faster R-CNN. Each model page lists its full task coverage, license, and specs.
2 of the 5 Microsoft vision models we track handle OCR: Florence-2 and TrOCR. Each model page lists its full task coverage, license, and specs.
Yes. All 5 Microsoft vision models we track publish downloadable weights you can self-host under their licenses (MIT and Custom).
We do not publish a Microsoft-only ranking, so pick on constraints rather than a label. 2 Microsoft models handle object detection; the most recent is Florence-2 (Jun 2025). Measured scores across every lab we test are on our object detection benchmark, linked at the bottom of this page. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.
We track 5 live Microsoft vision models. The most recent addition is Florence-2, released Jun 2025.
Yes. 1 of the 5 Microsoft models run live in the Roboflow Playground. Upload your own image, run several models on it at once, and compare the outputs side by side. No setup and no account required.
This page lists all 5 Microsoft vision models in the Roboflow Playground catalog, all of them open-weight and free to self-host. They cover object detection, OCR, and captioning, among other tasks. 1 of them run live in the Roboflow Playground on your own images. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.