groundingdino-py
open-set object detector
Decision gist · record as of 2026-08-14
No. While the model itself is capable, the PyPI package is problematic: it has been dormant since December 2023 with no maintenance, installation requires compiling native code with high friction, and the fact sheet lists zero runtime dependencies despite being a deep learning model—suggesting the package metadata is incomplete or the PyPI distribution is not the intended installation path. The GitHub repository is the canonical source; install from there directly if you need this model.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA_HOME environment variable to be set for GPU compilation; falls back to CPU-only mode if CUDA is unavailable.
- Model weights must be downloaded separately.
- Compilation of native code required during installation.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache License 2.0, a permissive license that allows commercial and private use with minimal restrictions, requiring only attribution and notice of modifications.
last release 2023-05-23 (1179 days) · last repo commit 2023-12-20 · 13 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 133,313 downloads/mo, #11,520 on PyPI
Alternatives
Verify before relying
# Clone and install from source
git clone https://github.com/IDEA-Research/GroundingDINO.git
cd GroundingDINO
pip install -e .
# Download model weights
mkdir weights
cd weights
wget https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth
# Basic inference (see demo/inference_on_a_image.py)
CUDA_VISIBLE_DEVICES=0 python demo/inference_on_a_image.py \
-c groundingdino/config/GroundingDINO_SwinT_OGC.py \
-p weights/groundingdino_swint_ogc.pth \
-i image.jpg- Whether the package works with current PyTorch versions and modern Python releases, given dormant maintenance status since late 2023
- Actual runtime dependencies and their versions, as the fact sheet lists zero runtime dependencies despite being a deep learning model
- Whether the PyPI package is actively maintained or if the GitHub repository is the canonical source
What it is and what it does
Grounding DINO combines vision and language understanding to detect objects in images by their natural language descriptions. Unlike traditional object detectors that recognize only pre-trained classes, it can identify any object you describe in text, making it an open-set detector. The model outputs bounding boxes for detected objects along with confidence scores for each word in your text prompt, allowing you to filter results by similarity threshold.
The package is designed for research and production use in computer vision tasks where you need flexible, language-driven object detection. It accepts image-text pairs as input and outputs up to 900 candidate boxes by default, each scored against all input words. The implementation supports both GPU and CPU inference, though GPU is strongly recommended for practical use. Model weights must be downloaded separately from the GitHub releases.
Use it for
- Automated dataset annotation: use natural language prompts to label objects across large image collections without manual annotation
- Image editing workflows: identify specific objects by description to segment or edit them in combination with tools like Stable Diffusion or SAM
- Visual search and retrieval: find objects matching text descriptions across image databases without retraining for new object types
- Accessibility tools: describe what you want to find in an image and get bounding boxes for screen readers or assistive systems
- Content moderation: detect problematic objects or scenes by describing them in text without maintaining separate classifiers
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
While the model itself is capable, the PyPI package is problematic: it has been dormant since December 2023 with no maintenance, installation requires compiling native code with high friction, and the fact sheet lists zero runtime dependencies despite being a deep learning model—suggesting the package metadata is incomplete or the PyPI distribution is not the intended installation path. The GitHub repository is the canonical source; install from there directly if you need this model.
Install
groundingdino-py on PyPI
Before you install
Installation requires compiling native code and has high friction. The package is dormant—last commit was 2023-12-20, over a year ago, with no active maintenance. Expect potential compatibility issues with newer Python or dependency versions.
Requires CUDA_HOME environment variable to be set for GPU compilation; falls back to CPU-only mode if CUDA is unavailable. Model weights must be downloaded separately. Compilation of native code required during installation.
License in practice
Licensed under Apache License 2.0, a permissive license that allows commercial and private use with minimal restrictions, requiring only attribution and notice of modifications.
Quickstart
# Clone and install from source
git clone https://github.com/IDEA-Research/GroundingDINO.git
cd GroundingDINO
pip install -e .
# Download model weights
mkdir weights
cd weights
wget https://github.com/IDEA-Research/GroundingDINO/releases/download/v0.1.0-alpha/groundingdino_swint_ogc.pth
# Basic inference (see demo/inference_on_a_image.py)
CUDA_VISIBLE_DEVICES=0 python demo/inference_on_a_image.py \
-c groundingdino/config/GroundingDINO_SwinT_OGC.py \
-p weights/groundingdino_swint_ogc.pth \
-i image.jpg
Verify before relying
- Whether the package works with current PyTorch versions and modern Python releases, given dormant maintenance status since late 2023
- Actual runtime dependencies and their versions, as the fact sheet lists zero runtime dependencies despite being a deep learning model
- Whether the PyPI package is actively maintained or if the GitHub repository is the canonical source
Package facts
| License | permissive license permissive |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Dormant 1,179 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 133,313 / month, #11,520 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: groundingdino-py-0.4.0.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “zero-shot object detection”
- groundingdino-pyGrounding DINO is an open-set object detector that identifies and…
- nixtlaPython SDK for accessing TimeGPT, a foundation model for time series…
- glinerGLiNER is a lightweight framework for named entity recognition that…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also mmdet · perceptron · ultralytics · yolov5 · lightly · rfdetr · effdet · segmentation-models-pytorch · speechbrain · sahi