$npx skillfedfor your agent

databricks-labs-dqx

Data Quality eXtended (DQX) is a Python library for data quality checks and data quality monitoring

With conditionsPyPI LibrariesReleased Aug 20267.5M downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — databricks_labs_dqx-0.16.0-py3-none-any.whl
v0.16.0 · released 2026-08-13 · Python >=3.10 · 5 runtime deps: databricks-labs-blueprint, databricks-sdk, pydantic, pyyaml, sqlalchemy

Yes, with conditions. DQX is actively maintained, has low install friction, and offers a comprehensive rule-based quality framework tailored to Databricks workloads. However, the license treatment is unclear—verify the actual license before production use. Also confirm that advanced features fit your Databricks setup and cost model. Not formally supported by Databricks with SLAs; issues are reviewed as time permits.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; databricks-sdk must be configured with valid Databricks workspace credentials.
  • Low friction: pure Python wheel with five runtime dependencies (databricks-sdk, pydantic, pyyaml, sqlalchemy, databricks-labs-blueprint).
  • Active maintenance with a release 1 day old and recent commits.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available. Verify the actual license terms before adopting in production, especially for commercial use.

last release 2026-08-13 (1 days) · last repo commit 2026-08-14 · 447 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 7,497,459 downloads/mo, #1,735 on PyPI

Verify before relying

pip install databricks-labs-dqx

from databricks_labs_dqx import DQX

dqx = DQX()
results = dqx.run_checks(df, checks=[...])
  • Exact scope and maturity level of the built-in checks across different categories (null, range, regex, referential, aggregate, geo, PII).
  • Performance characteristics and scalability limits for very large datasets or high-frequency streaming workloads.
  • Whether LLM-driven rule generation and anomaly detection require additional Databricks services or incur extra costs.
  • Compatibility and integration details with Databricks Unity Catalog, Volumes, and Lakebase (PostgreSQL) backends.
Same gist for agents: .md · .json

What it is and what it does

DQX is a data quality framework for PySpark workloads on Databricks that lets you define, run, and monitor quality checks on both batch and streaming DataFrames. It supports built-in checks (null, range, regex, referential, aggregate, geo, PII) that you can apply row-level or column/dataset-level, plus custom check functions. You can define checks as code or declaratively in YAML/JSON, mark failures as warnings or errors, and route invalid data to quarantine, drop, or mark operations. The library includes data profiling, automatic rule generation from existing data, ML-based row anomaly detection, and integration with Databricks data contracts for schema validation.

Results are persisted to Delta tables with built-in aggregate metrics and a Lakeview dashboard for tracking quality over time. You can trigger Slack, Teams, or webhook alerts when metrics cross thresholds, and optionally fail pipelines on quality violations. DQX Studio provides a browser-based no-code UI for authoring and monitoring rules as a Databricks App. The framework applies the same API to batch DataFrames and Spark Structured Streaming (including DLT pipelines).

Use it for

  • Validate incoming data against business rules before loading into a data lake.
  • Monitor data quality metrics on streaming pipelines and alert teams when error rates exceed thresholds.
  • Generate quality rules automatically from existing datasets and refine them with LLM suggestions.
  • Enforce schema and referential integrity checks as part of a Databricks data contract.
  • Detect anomalous rows in large datasets and get explanations for investigation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

DQX is actively maintained, has low install friction, and offers a comprehensive rule-based quality framework tailored to Databricks workloads. However, the license treatment is unclear—verify the actual license before production use. Also confirm that advanced features fit your Databricks setup and cost model. Not formally supported by Databricks with SLAs; issues are reviewed as time permits.

Install

databricks-labs-dqx on PyPI

Before you install

Low friction: pure Python wheel with five runtime dependencies (databricks-sdk, pydantic, pyyaml, sqlalchemy, databricks-labs-blueprint). Active maintenance with a release 1 day old and recent commits. Requires Python 3.10+.

Requires Python 3.10 or later; databricks-sdk must be configured with valid Databricks workspace credentials.

License in practice

License treatment is unclear—no SPDX identifier or raw license text is available. Verify the actual license terms before adopting in production, especially for commercial use.

Quickstart

pip install databricks-labs-dqx

from databricks_labs_dqx import DQX

dqx = DQX()
results = dqx.run_checks(df, checks=[...])

Verify before relying

  • Exact scope and maturity level of the built-in checks across different categories (null, range, regex, referential, aggregate, geo, PII).
  • Performance characteristics and scalability limits for very large datasets or high-frequency streaming workloads.
  • Whether LLM-driven rule generation and anomaly detection require additional Databricks services or incur extra costs.
  • Compatibility and integration details with Databricks Unity Catalog, Volumes, and Lakebase (PostgreSQL) backends.

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
5 packages
databricks-labs-blueprintdatabricks-sdkpydanticpyyamlsqlalchemy
MaintenanceActively maintained 1 days since the last release
Last repo commit
First released
Downloads7,497,459 / month, #1,735 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: ConsoleFramework :: PytestIntended Audience :: DevelopersIntended Audience :: System AdministratorsLicense :: Other/Proprietary LicenseOperating System :: MacOSOperating System :: Microsoft :: WindowsProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: Implementation :: CPythonTopic :: Software Development :: LibrariesTopic :: Utilities

Evidence: databricks_labs_dqx-0.16.0-py3-none-any.whl

Tags

Capabilities
pyspark data quality validationdatabricks data quality checksrule-based data profilingstreaming dataframe validationdata quality monitoring frameworkdata contracts validationautomated data validation at scale
Topics
databricksdata-qualitystreaming
PyPI keywords
Databricks

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “databricks data quality checks”

  • databricks-labs-dqxDQX provides rule-based data quality checking for PySpark DataFrames…
  • dlt-metaDLT-META is a metadata-driven framework that automates bronze and…
  • datacontract-cliA CLI tool for defining, validating, and testing data contracts using…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also dbldatagen · spark-expectations · cuallee · pydeequ · dbt-databricks · databricks-labs-remorph · databricks-labs-lsql · databricks-test · pyddq · panzi-json-logic