Roboflow

IDEA Research Vision Models

Compare both IDEA Research vision models we track, all of them open-weight.

2 models · 2 open-weight · Updated Jan 2024

About IDEA Research Vision Models

Every IDEA Research vision model in our catalog is open-weight. Both publish downloadable weights, so you can self-host under their own licenses instead of paying per request. The open-weight side runs on Apache 2.0 terms. The newest additions are Grounded SAM (Jan 2024) and Grounding DINO (Mar 2023).

Between them, the IDEA Research models we list cover 4 distinct vision tasks. The widest coverage is object detection (2 models), open vocabulary object detection (2), and vision language (1). For object detection, start with Grounded SAM or Grounding DINO.

On the open-weight tier, Grounding DINO runs at 172M-341M parameters. Both carry a permissive license, so commercial use is straightforward.

With every model here open-weight, the real trade is size against generality: the smaller checkpoints fine-tune and deploy cheaply on your own hardware, while the larger ones cover more ground out of the box. Compare their licenses, sizes, prices, and release dates in the tables below, then open any model page for the full specification.

Which IDEA Research Model Should You Use?

What each of the 2 IDEA Research vision models in our catalog is built for, and how you run it.

Grounded SAM
Best for: Zero Shot Segmentation and Open Vocabulary Object Detection. Open weights: Apache 2.0 license. Self-host it or run it here.
Grounding DINO
Best for: Open Vocabulary Object Detection and Object Detection. Open weights: Apache 2.0 license, 172M-341M parameters. Self-host it or run it here. Small enough for on-device and edge deployment.

Open-Source IDEA Research Models

2 models with downloadable weights you can self-host under their licenses (Apache 2.0).

IDEA Research
Grounded SAM
Grounded SAM is an open-vocabulary image segmentation model developed by IDEA Research, released in January 2024 under the Apache 2.0 license. It combines Grounding DINO, a zero-shot open-vocabulary object detector, with the Segment Anything Model to produce precise segmentation masks for objects identified through free-form text prompts. The two models are used sequentially: Grounding DINO localizes objects from a text query, and SAM generates the corresponding segmentation masks.Grounded SAM enables zero-shot instance segmentation without task-specific training data, making it applicable to domains where labeled segmentation data is scarce. It supports arbitrary text queries and can segment objects not represented in standard training sets. The model is commonly used in automated labeling pipelines, robotic perception, and domain-specific vision applications requiring open-vocabulary segmentation.
IDEA Research
Grounding DINO
Grounding DINO is an open-vocabulary object detection model developed by IDEA Research, released in March 2023 under the Apache 2.0 license. It extends the DINO transformer-based detector with grounded pre-training, enabling it to detect arbitrary objects described by free-form text queries rather than a fixed set of predefined categories. The model integrates a text encoder with a visual backbone through a feature fusion module that aligns language and visual representations at multiple scales.Grounding DINO achieves strong zero-shot detection performance on COCO, LVIS, and ODinW benchmarks, and supports referring expression comprehension tasks. It is widely used as a foundation for open-vocabulary detection pipelines and as the detection backbone in systems such as Grounded-SAM. The model is particularly suited for applications requiring flexible, text-driven object localization across diverse domains.

Frequently Asked Questions About IDEA Research Vision Models

Which IDEA Research models can do object detection?

2 of the 2 IDEA Research vision models we track handle object detection: Grounded SAM and Grounding DINO. Each model page lists its full task coverage, license, and specs.

Which IDEA Research models can do open vocabulary object detection?

2 of the 2 IDEA Research vision models we track handle open vocabulary object detection: Grounded SAM and Grounding DINO. Each model page lists its full task coverage, license, and specs.

Are IDEA Research vision models open source?

Yes. Both IDEA Research vision models we track publish downloadable weights you can self-host under their licenses (Apache 2.0).

What is the best IDEA Research model for object detection?

We do not publish an IDEA Research-only ranking, so pick on constraints rather than a label. 2 IDEA Research models handle object detection; the most recent is Grounded SAM (Jan 2024). Measured scores across every lab we test are on our object detection benchmark, linked at the bottom of this page. For fixed categories in production, a model fine-tuned on your own data typically beats any general-purpose model.

How many IDEA Research vision models are on Roboflow Playground?

We track 2 live IDEA Research vision models. The most recent addition is Grounded SAM, released Jan 2024.

This page lists both IDEA Research vision models in the Roboflow Playground catalog, all of them open-weight and free to self-host. They cover object detection, open vocabulary object detection, and vision language, among other tasks. Compare licenses, parameters, prices, and release dates side by side, or open any model page for full details.