{"categories":[{"label":"Testing","url":"https://skillfed.io/packages/category/software-development-testing/3"}],"enrichment":{"capability":"Generates synthetic data at scale within Databricks using Spark, supporting repeatable, predictable datasets for testing, benchmarking, and demos across all Spark SQL primitive types.","skillfed_tags":["databricks","synthetic-data","spark"],"use_cases":["Generate large test datasets for validating ETL pipelines and data transformations in Databricks without exposing real data.","Create repeatable synthetic data for demos and benchmarking to measure cluster performance under controlled, reproducible conditions.","Produce consistent multi-table datasets with matching foreign keys for testing join, merge, and Change Data Capture scenarios.","Build feature arrays for machine learning model testing using weighted distributions and custom column expressions.","Populate Delta Live Tables pipelines with synthetic data sources for end-to-end pipeline validation."],"what_it_does":"dbldatagen is a PySpark library for generating synthetic data within Databricks environments. It lets you define data generation specifications in code to produce repeatable, predictable datasets at scale\u2014up to billions of rows\u2014for testing, benchmarking, and demos. You can generate data conforming to existing schemas or create ad-hoc datasets, with support for all Spark SQL primitive types, date/timestamp ranges, discrete values, weighted distributions, arrays, and SQL expressions. It includes a plugin mechanism for third-party libraries and integrates with Delta Live Tables pipelines.\n\nThe library has no dependencies on packages outside the Databricks runtime, making it lightweight to install and use. It supports generating data with consistency between primary and foreign keys for join and merge scenarios, and can produce code from existing schemas. You can access generated data from Scala, R, and other languages via views, and persist or export results to external storage.","worth_installing":"Yes, if you work within Databricks and need synthetic data generation. The library is actively maintained, has no external dependencies, and integrates seamlessly with Spark. However, verify the 'Databricks License' terms for your use case before installing, particularly if you plan to use it outside Databricks or in proprietary applications."},"id":"dbldatagen","links":{"html":"https://skillfed.io/packages/dbldatagen","md":"https://skillfed.io/packages/dbldatagen.md","pypi":"https://pypi.org/project/dbldatagen/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2024-07-26","license_spdx":null,"license_treatment":"unclear","name":"dbldatagen","python_support":"supports_current","summary":"Databricks Labs -  PySpark Synthetic Data Generator"},"popularity":{"monthly_downloads":357181,"position":7274,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.4.0.post1"}
