dbldatagen
Databricks Labs - PySpark Synthetic Data Generator
Decision gist · record as of 2026-08-14
Yes, if you work within Databricks and need synthetic data generation. The library is actively maintained, has no external dependencies, and integrates seamlessly with Spark. However, verify the 'Databricks License' terms for your use case before installing, particularly if you plan to use it outside Databricks or in proprietary applications.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires PySpark 3.1.2 and Python 3.8 or later; designed for Databricks runtime 10.4 LTS and later.
- Databricks runtime 13.2 or later recommended for full Unity Catalog support.
- Low install friction with no runtime dependencies outside Databricks.
License · maintenance · safety
Databricks License (unclear) — License treatment is unclear—listed as 'Databricks License' with no SPDX identifier. Review the actual license terms before use in proprietary or redistributed contexts.
last release 2024-07-26 (749 days) · last repo commit 2026-08-07 · 488 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 357,181 downloads/mo, #7,274 on PyPI
Alternatives
Verify before relying
import dbldatagen as dg
df = dg.Datasets(spark, "basic/user").get(rows=1000_000).build()
num_rows = df.count()- Whether the 'Databricks License' permits use outside Databricks environments or in closed-source applications.
- Performance characteristics and scalability limits when generating billions of rows on specific cluster configurations.
- Plugin mechanism details and compatibility with third-party libraries like Faker beyond what the description states.
What it is and what it does
dbldatagen is a PySpark library for generating synthetic data within Databricks environments. It lets you define data generation specifications in code to produce repeatable, predictable datasets at scale—up to billions of rows—for testing, benchmarking, and demos. You can generate data conforming to existing schemas or create ad-hoc datasets, with support for all Spark SQL primitive types, date/timestamp ranges, discrete values, weighted distributions, arrays, and SQL expressions. It includes a plugin mechanism for third-party libraries and integrates with Delta Live Tables pipelines.
The library has no dependencies on packages outside the Databricks runtime, making it lightweight to install and use. It supports generating data with consistency between primary and foreign keys for join and merge scenarios, and can produce code from existing schemas. You can access generated data from Scala, R, and other languages via views, and persist or export results to external storage.
Use it for
- Generate large test datasets for validating ETL pipelines and data transformations in Databricks without exposing real data.
- Create repeatable synthetic data for demos and benchmarking to measure cluster performance under controlled, reproducible conditions.
- Produce consistent multi-table datasets with matching foreign keys for testing join, merge, and Change Data Capture scenarios.
- Build feature arrays for machine learning model testing using weighted distributions and custom column expressions.
- Populate Delta Live Tables pipelines with synthetic data sources for end-to-end pipeline validation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work within Databricks and need synthetic data generation.
The library is actively maintained, has no external dependencies, and integrates seamlessly with Spark. However, verify the 'Databricks License' terms for your use case before installing, particularly if you plan to use it outside Databricks or in proprietary applications.
Install
dbldatagen on PyPI
Before you install
Low install friction with no runtime dependencies outside Databricks. Actively maintained with recent commits; last release 749 days ago indicates a stable, mature project rather than rapid iteration.
Requires PySpark 3.1.2 and Python 3.8 or later; designed for Databricks runtime 10.4 LTS and later. Databricks runtime 13.2 or later recommended for full Unity Catalog support.
License in practice
License treatment is unclear—listed as 'Databricks License' with no SPDX identifier. Review the actual license terms before use in proprietary or redistributed contexts.
Quickstart
import dbldatagen as dg
df = dg.Datasets(spark, "basic/user").get(rows=1000_000).build()
num_rows = df.count()
Verify before relying
- Whether the 'Databricks License' permits use outside Databricks environments or in closed-source applications.
- Performance characteristics and scalability limits when generating billions of rows on specific cluster configurations.
- Plugin mechanism details and compatibility with third-party libraries like Faker beyond what the description states.
Package facts
| License | Databricks License unclear |
| Python support | Supports the current Python release >=3.8.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 749 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 357,181 / month, #7,274 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersIntended Audience :: System AdministratorsOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Software Development :: Testing |
Evidence: dbldatagen-0.4.0.post1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “synthetic data generation spark”
- dbldatagenGenerates synthetic data at scale within Databricks using Spark,…
- pyspark-data-sourcesProvides custom Apache Spark data sources using the Python Data…
- data-designer-engineExecution engine for the NeMo Data Designer synthetic data generation…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also databricks-labs-dqx · databricks-test · dbl-tempo · databricks-labs-remorph · dbt-databricks · dlt-meta · pyspark-data-sources · dbx · dbl-discoverx · hyperleaup