{"categories":[{"label":"Libraries","url":"https://skillfed.io/packages/category/software-development-libraries/2"},{"label":"Utilities","url":"https://skillfed.io/packages/category/utilities/2"}],"enrichment":{"capability":"DQX provides rule-based data quality checking for PySpark DataFrames on Databricks, supporting batch and streaming workloads with built-in checks, custom rules, and automated quality monitoring.","skillfed_tags":["databricks","data-quality","streaming"],"use_cases":["Validate incoming data against business rules before loading into a data lake.","Monitor data quality metrics on streaming pipelines and alert teams when error rates exceed thresholds.","Generate quality rules automatically from existing datasets and refine them with LLM suggestions.","Enforce schema and referential integrity checks as part of a Databricks data contract.","Detect anomalous rows in large datasets and get explanations for investigation."],"what_it_does":"DQX is a data quality framework for PySpark workloads on Databricks that lets you define, run, and monitor quality checks on both batch and streaming DataFrames. It supports built-in checks (null, range, regex, referential, aggregate, geo, PII) that you can apply row-level or column/dataset-level, plus custom check functions. You can define checks as code or declaratively in YAML/JSON, mark failures as warnings or errors, and route invalid data to quarantine, drop, or mark operations. The library includes data profiling, automatic rule generation from existing data, ML-based row anomaly detection, and integration with Databricks data contracts for schema validation.\n\nResults are persisted to Delta tables with built-in aggregate metrics and a Lakeview dashboard for tracking quality over time. You can trigger Slack, Teams, or webhook alerts when metrics cross thresholds, and optionally fail pipelines on quality violations. DQX Studio provides a browser-based no-code UI for authoring and monitoring rules as a Databricks App. The framework applies the same API to batch DataFrames and Spark Structured Streaming (including DLT pipelines).","worth_installing":"Yes, with conditions. DQX is actively maintained, has low install friction, and offers a comprehensive rule-based quality framework tailored to Databricks workloads. However, the license treatment is unclear\u2014verify the actual license before production use. Also confirm that advanced features fit your Databricks setup and cost model. Not formally supported by Databricks with SLAs; issues are reviewed as time permits."},"id":"databricks-labs-dqx","links":{"html":"https://skillfed.io/packages/databricks-labs-dqx","md":"https://skillfed.io/packages/databricks-labs-dqx.md","pypi":"https://pypi.org/project/databricks-labs-dqx/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-13","license_spdx":null,"license_treatment":"unclear","name":"databricks-labs-dqx","python_support":"supports_current","summary":"Data Quality eXtended (DQX) is a Python library for data quality checks and data quality monitoring"},"popularity":{"monthly_downloads":7497459,"position":1735,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"0.16.0"}
