$npx skillfedfor your agent

data-designer

General framework for synthetic data generation

Worth itPyPI Software DevelopmentReleased Aug 2026316.5K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — data_designer-0.9.1-py3-none-any.whl
v0.9.1 · released 2026-08-11 · Python >=3.10 · 15 runtime deps: data-designer-config, data-designer-engine, huggingface-hub, opentelemetry-api, opentelemetry-exporter-prometheus, opentelemetry-sdk, packaging, pandas

Yes. Data Designer is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and addresses a real need for controlled synthetic data generation beyond simple prompting. It's suitable for developers building datasets for ML training, testing, or privacy-preserving development. The async engine and multi-provider support are practical for production workflows. No known security vulnerabilities. Start with a preview to test your schema before committing to full generation.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • Requires at least one API key set (NVIDIA_API_KEY, OPENAI_API_KEY, or OPENROUTER_API_KEY) to generate LLM-based columns; statistical samplers work without external APIs.
  • Low friction installation as a pure-Python wheel.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes.

last release 2026-08-11 (3 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 316,497 downloads/mo, #7,676 on PyPI

Verify before relying

pip install data-designer

import data_designer.config as dd
from data_designer.interface import DataDesigner

data_designer = DataDesigner()
config_builder = dd.DataDesignerConfigBuilder()
config_builder.add_column(
    dd.SamplerColumnConfig(
        name="product_category",
        sampler_type=dd.SamplerType.CATEGORY,
        params=dd.CategorySamplerParams(
            values=["Electronics", "Clothing", "Home & Kitchen", "Books"],
        ),
    )
)
preview = data_designer.preview(config_builder=config_builder)
  • Actual performance improvement from the async engine on typical workloads and how to measure it
  • Whether the agent skill works reliably with coding agents other than Claude Code
  • Telemetry data collection scope and frequency beyond model names and token counts
Same gist for agents: .md · .json

What it is and what it does

Data Designer is a framework for building synthetic datasets that go beyond simple LLM prompting. It lets you define columns using statistical samplers (category, numeric distributions), LLM generators, or seed data, then orchestrates generation with dependency-aware field relationships. The package includes validators (Python, SQL, custom, and remote LLM-as-judge) to assess output quality and a preview mode to test configurations before full-scale runs.

The library runs on an async, cell-level engine that overlaps independent column generation and adapts concurrency per provider and model. It integrates with multiple LLM providers (NVIDIA Build, OpenAI, OpenRouter) and includes OpenTelemetry instrumentation for monitoring. Configuration is built programmatically via a builder API and can also be managed through CLI commands. Telemetry is enabled by default but can be disabled via environment variable.

Use it for

  • Generate diverse product review datasets with realistic correlations between category, rating, and review text for training classification models.
  • Create synthetic customer records with demographic attributes, purchase history, and behavioral patterns for privacy-preserving testing and development.
  • Build test datasets for data pipelines and analytics by sampling from statistical distributions and validating output against SQL or Python rules.
  • Augment small seed datasets by generating synthetic variations while maintaining statistical properties and field dependencies.
  • Evaluate LLM quality on domain-specific tasks using LLM-as-judge validators to score generated outputs before production use.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Data Designer is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and addresses a real need for controlled synthetic data generation beyond simple prompting. It's suitable for developers building datasets for ML training, testing, or privacy-preserving development. The async engine and multi-provider support are practical for production workflows. No known security vulnerabilities. Start with a preview to test your schema before committing to full generation.

Install

data-designer on PyPI

Before you install

Low friction installation as a pure-Python wheel. Active maintenance with a release 3 days old. Depends on 15 runtime packages including pandas, pydantic, and OpenTelemetry; most are common data and ML infrastructure libraries.

Requires Python 3.10 or later. Requires at least one API key set (NVIDIA_API_KEY, OPENAI_API_KEY, or OPENROUTER_API_KEY) to generate LLM-based columns; statistical samplers work without external APIs.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes.

Quickstart

pip install data-designer

import data_designer.config as dd
from data_designer.interface import DataDesigner

data_designer = DataDesigner()
config_builder = dd.DataDesignerConfigBuilder()
config_builder.add_column(
    dd.SamplerColumnConfig(
        name="product_category",
        sampler_type=dd.SamplerType.CATEGORY,
        params=dd.CategorySamplerParams(
            values=["Electronics", "Clothing", "Home & Kitchen", "Books"],
        ),
    )
)
preview = data_designer.preview(config_builder=config_builder)

Verify before relying

  • Actual performance improvement from the async engine on typical workloads and how to measure it
  • Whether the agent skill works reliably with coding agents other than Claude Code
  • Telemetry data collection scope and frequency beyond model names and token counts

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
15 packages
data-designer-configdata-designer-enginehuggingface-hubopentelemetry-apiopentelemetry-exporter-prometheusopentelemetry-sdkpackagingpandasprometheus-clientprompt-toolkitpyarrowpydanticpyyamlrichtyper
MaintenanceActively maintained 3 days since the last release
First released
Downloads316,497 / month, #7,676 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development

Evidence: data_designer-0.9.1-py3-none-any.whl

Tags

Capabilities
synthetic data generationllm-based data synthesisdataset generation frameworksynthetic data with validationstatistical data samplingseed data augmentationquality-controlled synthetic datasets
Topics
synthetic-datallm-generationdata-validation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “seed data augmentation”

  • data-designerGenerates high-quality synthetic datasets from scratch or seed data,…
  • pytorch-seedProvides reproducible random number generation for PyTorch by seeding…
  • batchgeneratorsv2Reimplemented data augmentation transforms for deep learning,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Software Development packages

typing-extensions Worth it
PyPI · Software Development · released Jul 2026

Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.

PSF-2.0pure Python · 3.9+
1.9Bdownloads / mo
numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
annotated-doc With conditions
PyPI · Software Development · released Jul 2026

Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.

MITpure Python · 3.9+
456.2Mdownloads / mo
typer Worth it
PyPI · Software Development · released Aug 2026

Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.

Install it if you are building CLIs in Python.

MITpure Python · 3.10+
369.3Mdownloads / mo
distlib With conditions
PyPI · Software Development · released Jun 2026

Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.

permissive licensepure Python
323.3Mdownloads / mo

See also data-designer-config · data-designer-engine · oil-reservoir-synthesizer · copulas · sdv · sdmetrics · deepecho · nemo-gym · ucimlrepo · nemo-evaluator

Further reading