$npx skillfedfor your agent

spark-expectations

This project helps us to run Data Quality Rules in flight while spark job is being run

With conditionsPyPI Quality AssuranceReleased Jun 2026489.9K downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — spark_expectations-2.10.1-py3-none-any.whl
v2.10.1 · released 2026-06-27 · Python <=3.13,>=3.9 · 4 runtime deps: pluggy, pyyaml, requests, sqlglot

Yes, if you run PySpark pipelines and need built-in data quality enforcement with automatic error quarantine and observability. The low install friction, active maintenance, and absence of known vulnerabilities make it safe to adopt. However, verify the license terms first—the metadata does not declare a license identifier—and confirm that PySpark is already in your environment, as it is not listed as an explicit runtime dependency.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 to 3.13 and a working PySpark environment; configuration via Constants class is mandatory before use.
  • Low install friction with a pure Python wheel and four lightweight runtime dependencies (pluggy, pyyaml, requests, sqlglot).
  • Maintenance is active with a recent release within 48 days.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before use in proprietary or restricted contexts.

last release 2026-06-27 (48 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 489,853 downloads/mo, #6,374 on PyPI

Verify before relying

pip install spark-expectations

from spark_expectations.config.user_config import Constants as user_config

se_user_conf = {
    user_config.se_notifications_enable_email: False,
    user_config.se_enable_obs_dq_report_result: True,
}
  • Whether the package requires PySpark as a system dependency (not listed in runtime deps but core to its function).
  • Actual license identifier and any redistribution or commercial-use restrictions.
  • Whether email and Slack notifications require additional setup beyond the configuration shown.
Same gist for agents: .md · .json

What it is and what it does

Spark Expectations is a PySpark-native data quality framework that enforces data integrity by validating records against configurable rules before they reach downstream consumers. It supports three rule types: row-level checks (e.g., column constraints), aggregate checks (e.g., sum or average thresholds), and query-based checks (e.g., custom SQL validation). Records that fail any rule are automatically quarantined to a separate error table with metadata about which rules failed, job context, and detailed statistics, while only passing records move downstream. This prevents bad data from propagating and eliminates the need for manual error detection or separate corrective processes.

The framework integrates observability features that generate reports from statistics tables and can send alert notifications via email or Slack using customizable Jinja templates. Configuration is centralized through a Constants-based user config dictionary, making it straightforward to enable notifications, set error thresholds, and control which detailed metrics are captured. It targets Python 3.9–3.13 and depends on pluggy, pyyaml, requests, and sqlglot for plugin support, configuration parsing, HTTP calls, and SQL parsing.

Use it for

  • Enforce data contracts in ETL pipelines by rejecting rows that violate business rules and routing them to an error table for analysis.
  • Monitor aggregate data quality (e.g., row counts, sums, averages) and trigger alerts when metrics fall outside acceptable ranges.
  • Prevent downstream teams from consuming malformed data by automatically filtering and quarantining failed records at the source.
  • Generate observability reports and email/Slack notifications summarizing data quality metrics and rule failures for each job run.
  • Implement corrective workflows by isolating error records with full metadata, enabling teams to identify root causes and reprocess data.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you run PySpark pipelines and need built-in data quality enforcement with automatic error quarantine and observability.

The low install friction, active maintenance, and absence of known vulnerabilities make it safe to adopt. However, verify the license terms first—the metadata does not declare a license identifier—and confirm that PySpark is already in your environment, as it is not listed as an explicit runtime dependency.

Install

spark-expectations on PyPI

Before you install

Low install friction with a pure Python wheel and four lightweight runtime dependencies (pluggy, pyyaml, requests, sqlglot). Maintenance is active with a recent release within 48 days.

Requires Python 3.9 to 3.13 and a working PySpark environment; configuration via Constants class is mandatory before use.

License in practice

License treatment is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before use in proprietary or restricted contexts.

Quickstart

pip install spark-expectations

from spark_expectations.config.user_config import Constants as user_config

se_user_conf = {
    user_config.se_notifications_enable_email: False,
    user_config.se_enable_obs_dq_report_result: True,
}

Verify before relying

  • Whether the package requires PySpark as a system dependency (not listed in runtime deps but core to its function).
  • Actual license identifier and any redistribution or commercial-use restrictions.
  • Whether email and Slack notifications require additional setup beyond the configuration shown.

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release <=3.13,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
pluggypyyamlrequestssqlglot
MaintenanceActively maintained 48 days since the last release
First released
Downloads489,853 / month, #6,374 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python

Evidence: spark_expectations-2.10.1-py3-none-any.whl

Tags

Capabilities
pyspark data quality validationspark data quality frameworkdata quality rules pysparkdata validation pipeline sparkerror quarantine data qualityspark dq monitoringdata integrity spark jobs
Topics
data-qualitypysparkpipeline-validation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “spark data quality framework”

  • spark-expectationsSpark Expectations is a data quality framework that validates PySpark…
  • pydeequPyDeequ is a Python API for Apache Spark-based data quality…
  • pyddqPyDDQ is a Python API for running data quality constraint checks on…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also databricks-labs-dqx · quinn · pyspark-test · great-expectations · sparkmeasure · great-expectations-experimental · expects · acryl-great-expectations · repartipy · pyspark-pandas