dvc
Git for data scientists - manage your code and data together
Decision gist · record as of 2026-08-14
Yes. DVC is actively maintained, has low install friction, and solves a real problem for ML teams managing data versioning and reproducible pipelines. The Apache-2.0 license is permissive. No known vulnerabilities. The large dependency tree (42 packages) is justified by its feature set, and optional storage backends let you add only what you need. Suitable for both individual data scientists and collaborative teams.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later.
- Optional cloud storage backends (s3, gs, azure, ssh) must be installed separately if using remote storage.
- Low install friction with a pure-Python wheel.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and note any modifications.
last release 2026-03-31 (136 days) · last repo commit 2026-08-10 · 15,818 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,627,704 downloads/mo, #2,964 on PyPI
Alternatives
Verify before relying
pip install dvc
import dvc.api
# Track data in DVC
# dvc add data.csv
# dvc push # to remote storage
# Or use CLI: dvc init, dvc add, dvc run, dvc repro- Whether DVC's experiment tracking requires external servers or works entirely locally as claimed
- Performance characteristics when working with very large datasets or complex pipelines
- Specific Git hosting platforms tested and officially supported for experiment collaboration
What it is and what it does
DVC is a version control system for data and machine learning models that integrates with Git. It lets you store data artifacts and models outside your repository while keeping metadata in Git, similar to Git-LFS but without requiring a server. You define reproducible pipelines (computational graphs) that specify how to build models from code, data, and commands, then run only the steps affected by your changes.
The tool supports local experiment tracking—you can prepare and run many experiments, compare their results by hyperparameters and metrics, and visualize performance plots. It works with multiple remote storage backends (S3, Azure, Google Cloud, SSH, etc.) for sharing and backing up your data cache. The package has 42 runtime dependencies including celery, networkx, hydra-core, and fsspec, enabling distributed task execution and flexible storage integration.
Use it for
- Version and share large datasets and trained models with team members using existing Git hosting (GitHub, GitLab)
- Build reproducible ML pipelines that automatically track which steps need to re-run when code or data changes
- Run and compare multiple experiments locally, filtering results by hyperparameters and metrics without external servers
- Integrate data pipelines with CI/CD workflows to automatically reproduce experiments on code changes
- Store data in cloud storage (S3, Azure, GCS) while keeping version metadata in Git for cost-effective large-scale projects
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
DVC is actively maintained, has low install friction, and solves a real problem for ML teams managing data versioning and reproducible pipelines. The Apache-2.0 license is permissive. No known vulnerabilities. The large dependency tree (42 packages) is justified by its feature set, and optional storage backends let you add only what you need. Suitable for both individual data scientists and collaborative teams.
Install
dvc on PyPI
Before you install
Low install friction with a pure-Python wheel. Actively maintained with recent commits and a large community (15818 stars). Supports Python 3.9 through 3.14. Optional storage-specific dependencies (s3, gs, azure, ssh) are available for cloud integration.
Requires Python 3.9 or later. Optional cloud storage backends (s3, gs, azure, ssh) must be installed separately if using remote storage.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and note any modifications.
Quickstart
pip install dvc
import dvc.api
# Track data in DVC
# dvc add data.csv
# dvc push # to remote storage
# Or use CLI: dvc init, dvc add, dvc run, dvc repro
Verify before relying
- Whether DVC's experiment tracking requires external servers or works entirely locally as claimed
- Performance characteristics when working with very large datasets or complex pipelines
- Specific Git hosting platforms tested and officially supported for experiment collaboration
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 42 packagesattrscelerycoloramaconfigobjdistrodpathdulwichdvc-datadvc-httpdvc-objectsdvc-renderdvc-studio-clientdvc-taskflatten-dictflufl.lockfsspecfuncygrandalfgtohydra-coreiterative-telemetrykombunetworkxomegaconfpackagingpathspecplatformdirspsutilpydotpygtrie |
| Maintenance | Actively maintained 136 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,627,704 / month, #2,964 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9 |
Evidence: dvc-3.67.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “experiment tracking git”
- dvcDVC is a command-line tool for versioning data and models alongside…
- dvcliveDVCLive logs machine learning metrics, parameters, and plots to…
- dvc-studio-clientClient library for posting experiment metrics, model registry data,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also dvc-azure · dvc-render · dvc-studio-client · dvclive · laboratory · dvc-ssh · dvc-objects · dvc-http · dvc-s3 · vcstool