{"categories":[{"label":"Software Development","url":"https://skillfed.io/packages/category/software-development/9"},{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/6"}],"enrichment":{"capability":"Generates high-quality synthetic datasets from scratch or seed data, with control over field relationships, statistical distributions, and built-in validation and quality scoring.","skillfed_tags":["synthetic-data","llm-generation","data-validation"],"use_cases":["Generate diverse product review datasets with realistic correlations between category, rating, and review text for training classification models.","Create synthetic customer records with demographic attributes, purchase history, and behavioral patterns for privacy-preserving testing and development.","Build test datasets for data pipelines and analytics by sampling from statistical distributions and validating output against SQL or Python rules.","Augment small seed datasets by generating synthetic variations while maintaining statistical properties and field dependencies.","Evaluate LLM quality on domain-specific tasks using LLM-as-judge validators to score generated outputs before production use."],"what_it_does":"Data Designer is a framework for building synthetic datasets that go beyond simple LLM prompting. It lets you define columns using statistical samplers (category, numeric distributions), LLM generators, or seed data, then orchestrates generation with dependency-aware field relationships. The package includes validators (Python, SQL, custom, and remote LLM-as-judge) to assess output quality and a preview mode to test configurations before full-scale runs.\n\nThe library runs on an async, cell-level engine that overlaps independent column generation and adapts concurrency per provider and model. It integrates with multiple LLM providers (NVIDIA Build, OpenAI, OpenRouter) and includes OpenTelemetry instrumentation for monitoring. Configuration is built programmatically via a builder API and can also be managed through CLI commands. Telemetry is enabled by default but can be disabled via environment variable.","worth_installing":"Yes. Data Designer is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and addresses a real need for controlled synthetic data generation beyond simple prompting. It's suitable for developers building datasets for ML training, testing, or privacy-preserving development. The async engine and multi-provider support are practical for production workflows. No known security vulnerabilities. Start with a preview to test your schema before committing to full generation."},"id":"data-designer","links":{"html":"https://skillfed.io/packages/data-designer","md":"https://skillfed.io/packages/data-designer.md","pypi":"https://pypi.org/project/data-designer/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-11","license_spdx":"Apache-2.0","license_treatment":"permissive","name":"data-designer","python_support":"supports_current","summary":"General framework for synthetic data generation"},"popularity":{"monthly_downloads":316497,"position":7676,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.9.1"}
