Compare 3 phrase grounding models and try them on your own image, free in the Roboflow Playground. 3 are open-weight, so you can self-host them for free under their licenses.
3 models · 3 open-weight · 3 free to try · prices synced Jul 22, 2026
3 models with downloadable weights you can self-host under their licenses (Apache 2.0, MIT, and GPL v3). All run live in the Playground.
Phrase grounding is the task of localizing the exact image region that a text phrase refers to, linking language to pixels. Given "the person holding the red umbrella", the model must resolve attributes and relations between objects, not just find any person, and return that specific box or mask; referring expression comprehension is the one-target special case. Models align phrase embeddings with region features learned from image-text pairs. Grounding enables natural-language selection in editing tools, human-robot instructions, and precise multimodal search. This page lists 3 phrase grounding models, including 3 open-weight options you can self-host; all of them run live in the Playground so you can test them on your own images.
It depends on your task and constraints. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model. Compare the phrase grounding models on this page and try them on your own image to see which fits.
Yes. 3 of the 3 phrase grounding models here are open-weight (for example Qwen3.6 35B A3B, Florence-2, and YOLO World), free to self-host under their licenses (Apache 2.0, MIT, and GPL v3).
Yes. You can run all 3 in the Roboflow Playground for free. Upload an image and compare the models' output side by side, no setup required.
Explore other vision task pages, compare models head to head, or browse the full model catalog.