ctgan
Create tabular synthetic data using a conditional GAN
Decision gist · record as of 2026-08-14
Yes, with conditions. Install CTGAN if you need to generate synthetic tabular data and accept its Pre-Alpha maturity and BUSL-1.1 licensing restrictions. The library is actively maintained, has low install friction, and no known vulnerabilities. However, verify that the Business Source License aligns with your use case (commercial use may be restricted), and consider the SDV wrapper library if you need higher-level APIs and preprocessing support. Not recommended for production systems requiring stability guarantees.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires torch (PyTorch) as a runtime dependency; data must be preprocessed with continuous values as floats, discrete as ints/strings, and no missing values.
- Low install friction with five runtime dependencies (numpy, pandas, torch, tqdm, rdt).
- Actively maintained with a recent release 182 days ago and ongoing repository activity.
License · maintenance · safety
BUSL-1.1 (unclear) — Licensed under BUSL-1.1 (Business Source License), which restricts commercial use until a specified date and requires careful review of your use case. This is not a permissive open-source license; verify compatibility with your project's licensing requirements.
last release 2026-02-13 (182 days) · last repo commit 2026-08-10 · 1,559 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 122,961 downloads/mo, #11,927 on PyPI
Alternatives
Verify before relying
from ctgan import CTGAN, load_demo
real_data = load_demo()
discrete_columns = ['workclass', 'education', 'occupation']
ctgan = CTGAN(epochs=10)
ctgan.fit(real_data, discrete_columns)
synthetic_data = ctgan.sample(1000)- Whether BUSL-1.1 licensing restrictions apply to your intended use case and timeline.
- Performance characteristics and memory requirements for large datasets or high-dimensional tables.
- Fidelity and privacy guarantees compared to alternatives in the SDV ecosystem.
What it is and what it does
CTGAN is a standalone deep learning library for generating synthetic tabular data using conditional GAN and TVAE models. It learns patterns from real data and produces synthetic records that preserve statistical properties while protecting privacy. The library is part of the Synthetic Data Vault ecosystem but can be used independently; it requires careful data preprocessing (no missing values, proper type handling) and depends on PyTorch for model training.
Typical workflows involve loading real data, specifying which columns are discrete, training a model for a set number of epochs, and sampling synthetic records. The library is in Pre-Alpha stage, meaning its API and behavior may change. It's most suitable for research, testing, and non-production synthetic data needs where you control the licensing implications.
Use it for
- Generate synthetic test datasets for machine learning model development without exposing real customer or sensitive data.
- Create privacy-preserving data samples for sharing with external teams or in research publications.
- Augment imbalanced training datasets by synthesizing additional records for underrepresented classes.
- Prototype data pipelines and validate ETL logic on realistic synthetic data before running against production.
- Evaluate data quality and statistical properties of synthetic data generation models in research contexts.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Install CTGAN if you need to generate synthetic tabular data and accept its Pre-Alpha maturity and BUSL-1.1 licensing restrictions. The library is actively maintained, has low install friction, and no known vulnerabilities. However, verify that the Business Source License aligns with your use case (commercial use may be restricted), and consider the SDV wrapper library if you need higher-level APIs and preprocessing support. Not recommended for production systems requiring stability guarantees.
Install
ctgan on PyPI
Before you install
Low install friction with five runtime dependencies (numpy, pandas, torch, tqdm, rdt). Actively maintained with a recent release 182 days ago and ongoing repository activity. Pre-Alpha status signals the library is still evolving; consider this for experimental or non-production workflows.
Requires torch (PyTorch) as a runtime dependency; data must be preprocessed with continuous values as floats, discrete as ints/strings, and no missing values.
License in practice
Licensed under BUSL-1.1 (Business Source License), which restricts commercial use until a specified date and requires careful review of your use case. This is not a permissive open-source license; verify compatibility with your project's licensing requirements.
Quickstart
from ctgan import CTGAN, load_demo
real_data = load_demo()
discrete_columns = ['workclass', 'education', 'occupation']
ctgan = CTGAN(epochs=10)
ctgan.fit(real_data, discrete_columns)
synthetic_data = ctgan.sample(1000)
Verify before relying
- Whether BUSL-1.1 licensing restrictions apply to your intended use case and timeline.
- Performance characteristics and memory requirements for large datasets or high-dimensional tables.
- Fidelity and privacy guarantees compared to alternatives in the SDV ecosystem.
Package facts
| License | BUSL-1.1 unclear |
| Python support | Supports the current Python release <3.15,>=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesnumpypandastorchtqdmrdt |
| Maintenance | Actively maintained 182 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 122,961 / month, #11,927 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 2 - Pre-AlphaIntended Audience :: DevelopersNatural Language :: EnglishProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: ctgan-0.12.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “synthetic tabular data generation”
- ctganCTGAN generates synthetic tabular data by training deep learning…
- sdvSDV generates synthetic tabular data by learning patterns from real…
- copulasCopulas models multivariate statistical distributions and generates…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also deepecho · sdv · copulas · sdmetrics · autogluon.tabular · rdt · autogluon · featuretools · albumentations · batchgeneratorsv2