sdv
Generate synthetic data for single table, multi table and sequential data
Decision gist · record as of 2026-08-14
Yes, with conditions. SDV is actively maintained, has no known vulnerabilities, and low install friction. It's a mature tool (Production/Stable status) for a real need—synthetic data generation with privacy controls. However, the BUSL-1.1 license treatment is unclear; verify the license terms match your use case (commercial, internal, or research) before committing to production deployment. If licensing is acceptable, it's a solid choice for tabular synthetic data work.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with a pure Python wheel.
- Active maintenance with a release 7 days ago and consistent development activity.
- Supports Python 3.9 through 3.14.
License · maintenance · safety
BUSL-1.1 (unclear) — Licensed under BUSL-1.1 (Business Source License). License treatment is marked unclear in the metadata—review the actual license terms before use, particularly for commercial applications.
last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 3,544 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 118,605 downloads/mo, #12,113 on PyPI
Alternatives
Verify before relying
from sdv.datasets.demo import download_demo
from sdv.single_table import GaussianCopulaSynthesizer
real_data, metadata = download_demo(modality='single_table', dataset_name='fake_hotel_guests')
synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.fit(data=real_data)
synthetic_data = synthesizer.sample(num_rows=500)- Whether BUSL-1.1 restrictions apply to your intended use case (commercial, internal, or research)
- Memory and compute requirements for large datasets or complex multi-table schemas
- Performance characteristics and scalability limits for production workloads
What it is and what it does
SDV is a Python library for generating synthetic tabular data that mimics real datasets while protecting sensitive information. It offers multiple machine learning models—from classical statistical methods like GaussianCopula to deep learning approaches like CTGAN—to learn and replicate patterns in your data. The library handles single tables, multiple connected tables, and sequential data, with built-in support for preprocessing, anonymization, and business rule constraints.
The typical workflow involves loading or preparing your real data with metadata, selecting a synthesizer model, fitting it to learn patterns, then sampling synthetic rows. SDV also provides evaluation tools to measure how well the synthetic data matches the real data's statistical properties and to visualize differences. It depends on a substantial stack including pandas, numpy, copulas, ctgan, deepecho, rdt, and sdmetrics for its core functionality.
Use it for
- Generate test datasets for development and QA without exposing real customer or sensitive data
- Create shareable datasets for research or collaboration while maintaining privacy compliance
- Augment small datasets with synthetic rows to improve machine learning model training
- Evaluate data quality and statistical fidelity between real and synthetic versions
- Prototype data pipelines and analytics on realistic synthetic data before deploying to production
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
SDV is actively maintained, has no known vulnerabilities, and low install friction. It's a mature tool (Production/Stable status) for a real need—synthetic data generation with privacy controls. However, the BUSL-1.1 license treatment is unclear; verify the license terms match your use case (commercial, internal, or research) before committing to production deployment. If licensing is acceptable, it's a solid choice for tabular synthetic data work.
Install
sdv on PyPI
Before you install
Low install friction with a pure Python wheel. Active maintenance with a release 7 days ago and consistent development activity. Supports Python 3.9 through 3.14.
License in practice
Licensed under BUSL-1.1 (Business Source License). License treatment is marked unclear in the metadata—review the actual license terms before use, particularly for commercial applications.
Quickstart
from sdv.datasets.demo import download_demo
from sdv.single_table import GaussianCopulaSynthesizer
real_data, metadata = download_demo(modality='single_table', dataset_name='fake_hotel_guests')
synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.fit(data=real_data)
synthetic_data = synthesizer.sample(num_rows=500)
Verify before relying
- Whether BUSL-1.1 restrictions apply to your intended use case (commercial, internal, or research)
- Memory and compute requirements for large datasets or complex multi-table schemas
- Performance characteristics and scalability limits for production workloads
Package facts
| License | BUSL-1.1 unclear |
| Python support | Supports the current Python release <3.15,>=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 14 packagesboto3botocorecloudpicklegraphviznumpypandastqdmcopulasctgandeepechordtsdmetricsplatformdirspyyaml |
| Maintenance | Actively maintained 7 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 118,605 / month, #12,113 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersNatural Language :: EnglishProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: sdv-1.38.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “privacy-preserving data generation”
- sdvSDV generates synthetic tabular data by learning patterns from real…
- ctganCTGAN generates synthetic tabular data by training deep learning…
- flwrFlower is a framework for building federated learning systems where…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also copulas · ctgan · featuretools · rdt · sdmetrics · deepecho · oil-reservoir-synthesizer · data-designer · dbldatagen · pystac-ext-table