$npx skillfedfor your agent

pydeequ

PyDeequ - Unit Tests for Data

Worth itPyPI Quality AssuranceReleased Jul 202616.5M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pydeequ-1.6.0-py3-none-any.whl
v1.6.0 · released 2026-07-08 · Python <4,>=3.9 · 2 runtime deps: numpy, pandas

Yes. PyDeequ is actively maintained, has no known vulnerabilities, installs with low friction, and solves a concrete problem for data engineers. It is well-suited if you need declarative, scalable data quality checks. Caveat: you must have a Spark environment set up separately; PyDeequ is a wrapper, not a standalone tool.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a Spark environment and SparkSession configured with the Deequ Maven coordinate; Spark is not installed as a runtime dependency of PyDeequ.
  • Low install friction with a pure-Python wheel.
  • Active maintenance with a recent release 37 days ago and last commit on 2026-07-21.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and liability disclaimers.

last release 2026-07-08 (37 days) · last repo commit 2026-07-21 · 826 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 16,519,526 downloads/mo, #1,148 on PyPI

Verify before relying

pip install pydeequ

import pydeequ
from pydeequ.analyzers import AnalysisRunner, Size

result = AnalysisRunner(spark).onData(df).addAnalyzer(Size()).run()
  • Whether Spark is automatically installed or must be separately configured for PyDeequ to function
  • Performance characteristics and scalability limits for typical dataset sizes
  • Compatibility with specific Spark versions beyond what the classifiers indicate
Same gist for agents: .md · .json

What it is and what it does

PyDeequ wraps Deequ, an AWS-built library on Apache Spark, to define and run data quality checks at scale. It lets you compute metrics on dataset columns (via Analyzers and Profiles), suggest validation constraints based on data patterns, verify datasets against those constraints, and persist quality metrics over time in a repository. The package is designed for teams building data pipelines who need to catch data quality issues before they propagate downstream—treating data validation like unit tests for code.

You define checks declaratively (e.g., "column X must be complete and unique"), run them across large DataFrames, and get back pass/fail results at both aggregate and row levels. The row-level output lets you quarantine problematic records. It integrates with numpy and pandas as runtime dependencies.

Use it for

  • Run data profiling on large datasets to compute completeness, uniqueness, and distribution metrics before loading into a data warehouse.
  • Define and verify data quality constraints on incoming data in a data pipeline to catch schema violations or unexpected nulls early.
  • Suggest validation rules automatically based on sample data, then persist and track constraint results over time to monitor data quality trends.
  • Identify and quarantine rows that fail specific quality checks (e.g., invalid email format, out-of-range values) for manual review or remediation.
  • Build data quality dashboards by storing verification results in a metrics repository and querying historical runs by tag and timestamp.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

PyDeequ is actively maintained, has no known vulnerabilities, installs with low friction, and solves a concrete problem for data engineers. It is well-suited if you need declarative, scalable data quality checks. Caveat: you must have a Spark environment set up separately; PyDeequ is a wrapper, not a standalone tool.

Install

pydeequ on PyPI

Before you install

Low install friction with a pure-Python wheel. Active maintenance with a recent release 37 days ago and last commit on 2026-07-21. Depends only on numpy and pandas at runtime.

Requires a Spark environment and SparkSession configured with the Deequ Maven coordinate; Spark is not installed as a runtime dependency of PyDeequ.

License in practice

Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions beyond attribution and liability disclaimers.

Quickstart

pip install pydeequ

import pydeequ
from pydeequ.analyzers import AnalysisRunner, Size

result = AnalysisRunner(spark).onData(df).addAnalyzer(Size()).run()

Verify before relying

  • Whether Spark is automatically installed or must be separately configured for PyDeequ to function
  • Performance characteristics and scalability limits for typical dataset sizes
  • Compatibility with specific Spark versions beyond what the classifiers indicate

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release <4,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
numpypandas
MaintenanceActively maintained 37 days since the last release
Last repo commit
First released
Downloads16,519,526 / month, #1,148 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9

Evidence: pydeequ-1.6.0-py3-none-any.whl

Tags

Capabilities
data quality testing sparkunit tests for datadata profiling validationconstraint verification large datasetsdata quality metrics computationdeequ python apidata validation framework
Topics
data-qualitysparkdata-validation
PyPI keywords
deequpydeequdata-engineeringdata-qualitydata-profilingdataqualitydataunittestdata-unit-testsdata-profilers

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “constraint verification large datasets”

  • pydeequPyDeequ is a Python API for Apache Spark-based data quality…
  • pyvscPyVSC generates randomized test stimulus and defines and collects…
  • PyTestArchPyTestArch lets you define and test architectural rules for Python…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also cuallee · pyddq · tdda · sparkmeasure · databricks-labs-dqx · pyspark-extension · spark-expectations · quinn