$npx skillfedfor your agent

sdv

Generate synthetic data for single table, multi table and sequential data

With conditionsPyPI Artificial IntelligenceReleased Aug 2026118.6K downloads / moBUSL-1.1Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — sdv-1.38.0-py3-none-any.whl
v1.38.0 · released 2026-08-07 · Python <3.15,>=3.9 · 14 runtime deps: boto3, botocore, cloudpickle, graphviz, numpy, pandas, tqdm, copulas

Yes, with conditions. SDV is actively maintained, has no known vulnerabilities, and low install friction. It's a mature tool (Production/Stable status) for a real need—synthetic data generation with privacy controls. However, the BUSL-1.1 license treatment is unclear; verify the license terms match your use case (commercial, internal, or research) before committing to production deployment. If licensing is acceptable, it's a solid choice for tabular synthetic data work.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with a pure Python wheel.
  • Active maintenance with a release 7 days ago and consistent development activity.
  • Supports Python 3.9 through 3.14.

License · maintenance · safety

BUSL-1.1 (unclear) — Licensed under BUSL-1.1 (Business Source License). License treatment is marked unclear in the metadata—review the actual license terms before use, particularly for commercial applications.

last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 3,544 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 118,605 downloads/mo, #12,113 on PyPI

Verify before relying

from sdv.datasets.demo import download_demo
from sdv.single_table import GaussianCopulaSynthesizer

real_data, metadata = download_demo(modality='single_table', dataset_name='fake_hotel_guests')
synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.fit(data=real_data)
synthetic_data = synthesizer.sample(num_rows=500)
  • Whether BUSL-1.1 restrictions apply to your intended use case (commercial, internal, or research)
  • Memory and compute requirements for large datasets or complex multi-table schemas
  • Performance characteristics and scalability limits for production workloads
Same gist for agents: .md · .json

What it is and what it does

SDV is a Python library for generating synthetic tabular data that mimics real datasets while protecting sensitive information. It offers multiple machine learning models—from classical statistical methods like GaussianCopula to deep learning approaches like CTGAN—to learn and replicate patterns in your data. The library handles single tables, multiple connected tables, and sequential data, with built-in support for preprocessing, anonymization, and business rule constraints.

The typical workflow involves loading or preparing your real data with metadata, selecting a synthesizer model, fitting it to learn patterns, then sampling synthetic rows. SDV also provides evaluation tools to measure how well the synthetic data matches the real data's statistical properties and to visualize differences. It depends on a substantial stack including pandas, numpy, copulas, ctgan, deepecho, rdt, and sdmetrics for its core functionality.

Use it for

  • Generate test datasets for development and QA without exposing real customer or sensitive data
  • Create shareable datasets for research or collaboration while maintaining privacy compliance
  • Augment small datasets with synthetic rows to improve machine learning model training
  • Evaluate data quality and statistical fidelity between real and synthetic versions
  • Prototype data pipelines and analytics on realistic synthetic data before deploying to production

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

SDV is actively maintained, has no known vulnerabilities, and low install friction. It's a mature tool (Production/Stable status) for a real need—synthetic data generation with privacy controls. However, the BUSL-1.1 license treatment is unclear; verify the license terms match your use case (commercial, internal, or research) before committing to production deployment. If licensing is acceptable, it's a solid choice for tabular synthetic data work.

Install

sdv on PyPI

Before you install

Low install friction with a pure Python wheel. Active maintenance with a release 7 days ago and consistent development activity. Supports Python 3.9 through 3.14.

License in practice

Licensed under BUSL-1.1 (Business Source License). License treatment is marked unclear in the metadata—review the actual license terms before use, particularly for commercial applications.

Quickstart

from sdv.datasets.demo import download_demo
from sdv.single_table import GaussianCopulaSynthesizer

real_data, metadata = download_demo(modality='single_table', dataset_name='fake_hotel_guests')
synthesizer = GaussianCopulaSynthesizer(metadata)
synthesizer.fit(data=real_data)
synthetic_data = synthesizer.sample(num_rows=500)

Verify before relying

  • Whether BUSL-1.1 restrictions apply to your intended use case (commercial, internal, or research)
  • Memory and compute requirements for large datasets or complex multi-table schemas
  • Performance characteristics and scalability limits for production workloads

Package facts

LicenseBUSL-1.1 unclear
Python supportSupports the current Python release <3.15,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
14 packages
boto3botocorecloudpicklegraphviznumpypandastqdmcopulasctgandeepechordtsdmetricsplatformdirspyyaml
MaintenanceActively maintained 7 days since the last release
Last repo commit
First released
Downloads118,605 / month, #12,113 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersNatural Language :: EnglishProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial Intelligence

Evidence: sdv-1.38.0-py3-none-any.whl

Tags

Capabilities
synthetic data generationtabular data synthesisprivacy-preserving data generationmachine learning data simulationanonymized dataset creationstatistical data replicationmulti-table synthetic data
Topics
synthetic-dataprivacy-preservingdata-generation
PyPI keywords
sdvsynthetic-datasynthetic-data-generationtimeseriessingle-tablemulti-table

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “privacy-preserving data generation”

  • sdvSDV generates synthetic tabular data by learning patterns from real…
  • ctganCTGAN generates synthetic tabular data by training deep learning…
  • flwrFlower is a framework for building federated learning systems where…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also copulas · ctgan · featuretools · rdt · sdmetrics · deepecho · oil-reservoir-synthesizer · data-designer · dbldatagen · pystac-ext-table