$npx skillfedfor your agent

dbldatagen

Databricks Labs - PySpark Synthetic Data Generator

With conditionsPyPI TestingReleased Jul 2024357.2K downloads / moDatabricks LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — dbldatagen-0.4.0.post1-py3-none-any.whl
v0.4.0.post1 · released 2024-07-26 · Python >=3.8.10

Yes, if you work within Databricks and need synthetic data generation. The library is actively maintained, has no external dependencies, and integrates seamlessly with Spark. However, verify the 'Databricks License' terms for your use case before installing, particularly if you plan to use it outside Databricks or in proprietary applications.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires PySpark 3.1.2 and Python 3.8 or later; designed for Databricks runtime 10.4 LTS and later.
  • Databricks runtime 13.2 or later recommended for full Unity Catalog support.
  • Low install friction with no runtime dependencies outside Databricks.

License · maintenance · safety

Databricks License (unclear) — License treatment is unclear—listed as 'Databricks License' with no SPDX identifier. Review the actual license terms before use in proprietary or redistributed contexts.

last release 2024-07-26 (749 days) · last repo commit 2026-08-07 · 488 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 357,181 downloads/mo, #7,274 on PyPI

Verify before relying

import dbldatagen as dg
df = dg.Datasets(spark, "basic/user").get(rows=1000_000).build()
num_rows = df.count()
  • Whether the 'Databricks License' permits use outside Databricks environments or in closed-source applications.
  • Performance characteristics and scalability limits when generating billions of rows on specific cluster configurations.
  • Plugin mechanism details and compatibility with third-party libraries like Faker beyond what the description states.
Same gist for agents: .md · .json

What it is and what it does

dbldatagen is a PySpark library for generating synthetic data within Databricks environments. It lets you define data generation specifications in code to produce repeatable, predictable datasets at scale—up to billions of rows—for testing, benchmarking, and demos. You can generate data conforming to existing schemas or create ad-hoc datasets, with support for all Spark SQL primitive types, date/timestamp ranges, discrete values, weighted distributions, arrays, and SQL expressions. It includes a plugin mechanism for third-party libraries and integrates with Delta Live Tables pipelines.

The library has no dependencies on packages outside the Databricks runtime, making it lightweight to install and use. It supports generating data with consistency between primary and foreign keys for join and merge scenarios, and can produce code from existing schemas. You can access generated data from Scala, R, and other languages via views, and persist or export results to external storage.

Use it for

  • Generate large test datasets for validating ETL pipelines and data transformations in Databricks without exposing real data.
  • Create repeatable synthetic data for demos and benchmarking to measure cluster performance under controlled, reproducible conditions.
  • Produce consistent multi-table datasets with matching foreign keys for testing join, merge, and Change Data Capture scenarios.
  • Build feature arrays for machine learning model testing using weighted distributions and custom column expressions.
  • Populate Delta Live Tables pipelines with synthetic data sources for end-to-end pipeline validation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you work within Databricks and need synthetic data generation.

The library is actively maintained, has no external dependencies, and integrates seamlessly with Spark. However, verify the 'Databricks License' terms for your use case before installing, particularly if you plan to use it outside Databricks or in proprietary applications.

Install

dbldatagen on PyPI

Before you install

Low install friction with no runtime dependencies outside Databricks. Actively maintained with recent commits; last release 749 days ago indicates a stable, mature project rather than rapid iteration.

Requires PySpark 3.1.2 and Python 3.8 or later; designed for Databricks runtime 10.4 LTS and later. Databricks runtime 13.2 or later recommended for full Unity Catalog support.

License in practice

License treatment is unclear—listed as 'Databricks License' with no SPDX identifier. Review the actual license terms before use in proprietary or redistributed contexts.

Quickstart

import dbldatagen as dg
df = dg.Datasets(spark, "basic/user").get(rows=1000_000).build()
num_rows = df.count()

Verify before relying

  • Whether the 'Databricks License' permits use outside Databricks environments or in closed-source applications.
  • Performance characteristics and scalability limits when generating billions of rows on specific cluster configurations.
  • Plugin mechanism details and compatibility with third-party libraries like Faker beyond what the description states.

Package facts

LicenseDatabricks License unclear
Python supportSupports the current Python release >=3.8.10
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 749 days since the last release
Last repo commit
First released
Downloads357,181 / month, #7,274 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Intended Audience :: DevelopersIntended Audience :: System AdministratorsOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Software Development :: Testing

Evidence: dbldatagen-0.4.0.post1-py3-none-any.whl

Tags

Capabilities
synthetic data generation sparkdatabricks test data generatorpyspark fake data creationbulk test dataset generationdatabricks synthetic tables
Topics
databrickssynthetic-dataspark

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “synthetic data generation spark”

  • dbldatagenGenerates synthetic data at scale within Databricks using Spark,…
  • pyspark-data-sourcesProvides custom Apache Spark data sources using the Python Data…
  • data-designer-engineExecution engine for the NeMo Data Designer synthetic data generation…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also databricks-labs-dqx · databricks-test · dbl-tempo · databricks-labs-remorph · dbt-databricks · dlt-meta · pyspark-data-sources · dbx · dbl-discoverx · hyperleaup