$npx skillfedfor your agent

pyddq

Python API for Drunken Data Quality

SkipPyPI Quality AssuranceReleased Feb 2020106.3K downloads / moApache License Version 2.0Source build

Decision gist · record as of 2026-08-14

sdist only — pyddq-5.0.0.tar.gz · builds from source
v5.0.0 · released 2020-02-28

No. The project is abandoned (last update 2020-02-28), requires complex pre-configuration of a PySpark environment with external jar files, and has received no maintenance. Modern alternatives and actively maintained data validation libraries are strongly preferred for any production or new development work.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires PySpark environment with DDQ jar file already added to the driver class path; cannot be installed via pip alone without separate Spark and jar setup.
  • Installation friction is high: the package requires a pre-configured PySpark environment with the DDQ jar file already available, and the project has been abandoned since its last release with no maintenance activity.

License · maintenance · safety

Apache License Version 2.0 (permissive) — Licensed under Apache License Version 2.0 (permissive), allowing commercial use and modification without restriction, though the abandoned status means no ongoing legal or security updates.

last release 2020-02-28 (2359 days) · last repo commit 2020-02-28 · 220 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 106,337 downloads/mo, #12,656 on PyPI

Verify before relying

# Requires PySpark with DDQ jar pre-configured
pyspark --driver-class-path drunken-data-quality_2.11-5.0.0.jar

from pyddq.core import Check

df = spark.createDataFrame([(1, "a"), (1, None), (3, "c")])
check = Check(df)
check.hasUniqueKey("_1", "_2").isNeverNull("_1").run()
  • Whether the package works with Spark versions beyond 2.2.x (last tested version per documentation)
  • Python version compatibility (listed as unspecified in metadata)
  • Whether jar dependencies remain available or accessible from original sources
Same gist for agents: .md · .json

What it is and what it does

PyDDQ is a Python wrapper around the Drunken Data Quality (DDQ) Scala library, enabling data quality validation on Spark DataFrames through a fluent constraint-checking API. It lets you define and run checks for row counts, unique keys, foreign keys, nullability, and custom SQL expressions, then report results to stdout or custom reporters in formats like Markdown or console output.

The package is designed for continuous data import pipelines where you need to assert data quality before processing. However, it requires a pre-configured PySpark environment with the DDQ jar file loaded via the driver class path—it cannot be installed as a standalone Python package. The project has been abandoned since its last release on 2020-02-28, making it unsuitable for modern production environments without significant verification and potential maintenance work.

Use it for

  • Validate imported data meets row count and uniqueness constraints before ETL processing begins.
  • Write automated quality tests that inspect constraint results programmatically and fail data loads on violations.
  • Generate Markdown or console reports on data quality checks across multiple Spark DataFrames in a single run.
  • Enforce referential integrity by checking foreign key relationships between Spark tables.
  • Test that required columns are never null and custom SQL expressions hold true across the dataset.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No.

The project is abandoned (last update 2020-02-28), requires complex pre-configuration of a PySpark environment with external jar files, and has received no maintenance. Modern alternatives and actively maintained data validation libraries are strongly preferred for any production or new development work.

Install

pyddq on PyPI

Before you install

Installation friction is high: the package requires a pre-configured PySpark environment with the DDQ jar file already available, and the project has been abandoned since its last release with no maintenance activity.

Requires PySpark environment with DDQ jar file already added to the driver class path; cannot be installed via pip alone without separate Spark and jar setup.

License in practice

Licensed under Apache License Version 2.0 (permissive), allowing commercial use and modification without restriction, though the abandoned status means no ongoing legal or security updates.

Quickstart

# Requires PySpark with DDQ jar pre-configured
pyspark --driver-class-path drunken-data-quality_2.11-5.0.0.jar

from pyddq.core import Check

df = spark.createDataFrame([(1, "a"), (1, None), (3, "c")])
check = Check(df)
check.hasUniqueKey("_1", "_2").isNeverNull("_1").run()

Verify before relying

  • Whether the package works with Spark versions beyond 2.2.x (last tested version per documentation)
  • Python version compatibility (listed as unspecified in metadata)
  • Whether jar dependencies remain available or accessible from original sources

Package facts

LicenseApache License Version 2.0 permissive
Python supportNot specified
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceAbandoned 2,359 days since the last release
Last repo commit
First released
Downloads106,337 / month, #12,656 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaProgramming Language :: Python

Evidence: pyddq-5.0.0.tar.gz

Tags

Capabilities
spark dataframe validationdata quality constraintspyspark data testingconstraint checking sparkdata integrity validationpyspark quality assurancespark schema validation
Topics
spark-integrationdata-validationabandoned

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “data quality constraints”

  • pyddqPyDDQ is a Python API for running data quality constraint checks on…
  • dataframelyDataframely validates the schema and content of Polars data frames…
  • tddatdda provides test-driven data analysis tools: reference testing for…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also pydeequ · cuallee · pyspark-extension · pyspark-test · dbldatagen · databricks-labs-dqx · pyspark-stubs · quinn · tdda