datacompy
Dataframe comparisons in Python
Decision gist · record as of 2026-08-14
Yes. DataComPy is actively maintained, has no known vulnerabilities, low install friction, and solves a concrete problem for data engineers and QA teams. The permissive Apache license and support for multiple backends (Pandas, Polars, Spark, Snowflake) make it a practical choice for DataFrame validation and comparison workflows. Install it if you need to compare tabular data programmatically or generate detailed difference reports.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low install friction with a pure-wheel distribution.
- Active maintenance with a recent release (15 days ago) and ongoing commits.
License · maintenance · safety
Apache Software License (permissive) — Apache Software License (permissive) places no significant restrictions on use, modification, or distribution in proprietary or open-source projects.
last release 2026-07-30 (15 days) · last repo commit 2026-08-14 · 658 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 3,874,187 downloads/mo, #2,469 on PyPI
Alternatives
Verify before relying
pip install datacompy
import pandas as pd
from datacompy import PandasCompare
df1 = pd.DataFrame({"id": [1, 2, 3], "val": [10, 20, 30]})
df2 = pd.DataFrame({"id": [1, 2, 3], "val": [10, 99, 30]})
compare = PandasCompare(df1, df2, join_columns="id")
print(compare.report())- Whether tolerance/matching configuration options are documented and how granular they are.
- Performance characteristics when comparing very large DataFrames across backends.
- Whether Spark and Snowflake backends require additional system dependencies beyond pip extras.
What it is and what it does
DataComPy is a DataFrame comparison tool that goes beyond simple equality checks by identifying and reporting row and column differences across Pandas, Polars, Spark, and Snowflake tables. It was originally designed as a Python equivalent to SAS's PROC COMPARE, offering structured output that can be rendered as text reports, HTML files, or JSON-serializable dictionaries for programmatic use.
The package accepts two DataFrames, specifies join columns for row matching, and produces detailed statistics on mismatches. It exposes a ReportData API that lets you access comparison results programmatically without relying on string parsing, making it suitable for dashboards, automated validation pipelines, and data quality checks. Version 1.0.4 is the current stable release; the v0.19.x line is no longer supported.
Use it for
- Validate that ETL transformations produce expected output by comparing input and output DataFrames.
- Detect data drift or corruption by comparing snapshots of production tables at different points in time.
- Automate regression testing for data pipelines by comparing test results against baseline DataFrames.
- Generate audit reports showing which rows and columns differ between two database exports or Snowflake tables.
- Build data quality dashboards that consume structured comparison results via the ReportData API.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
DataComPy is actively maintained, has no known vulnerabilities, low install friction, and solves a concrete problem for data engineers and QA teams. The permissive Apache license and support for multiple backends (Pandas, Polars, Spark, Snowflake) make it a practical choice for DataFrame validation and comparison workflows. Install it if you need to compare tabular data programmatically or generate detailed difference reports.
Install
datacompy on PyPI
Before you install
Low install friction with a pure-wheel distribution. Active maintenance with a recent release (15 days ago) and ongoing commits. Runtime dependencies are well-established data libraries (pandas, polars, numpy, jinja2, ordered-set); optional extras for Spark and Snowflake are available but not required for core use.
Requires Python 3.10 or later.
License in practice
Apache Software License (permissive) places no significant restrictions on use, modification, or distribution in proprietary or open-source projects.
Quickstart
pip install datacompy
import pandas as pd
from datacompy import PandasCompare
df1 = pd.DataFrame({"id": [1, 2, 3], "val": [10, 20, 30]})
df2 = pd.DataFrame({"id": [1, 2, 3], "val": [10, 99, 30]})
compare = PandasCompare(df1, df2, join_columns="id")
print(compare.report())
Verify before relying
- Whether tolerance/matching configuration options are documented and how granular they are.
- Performance characteristics when comparing very large DataFrames across backends.
- Whether Spark and Snowflake backends require additional system dependencies beyond pip extras.
Package facts
| License | Apache Software License permissive |
| Python support | Supports the current Python release >=3.10.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesjinja2numpyordered-setpandaspolars |
| Maintenance | Actively maintained 15 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 3,874,187 / month, #2,469 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: datacompy-1.0.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “dataframe comparison”
- datacompyDataComPy compares two DataFrames across Pandas, Polars, Spark, and…
- pyspark-testProvides a testing utility to assert equality between two PySpark…
- chispaProvides assertion methods for testing PySpark code with descriptive…
Give your agent the search over MCP, or paste the wish link into any chat.
More Quality Assurance packages
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.
Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.
Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.
pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.
Install it if your test suite takes long enough that parallelization would save meaningful time.
Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.
Install it if you work with CloudFormation templates.
See also pyspark-test · dataframe-api-compat · recursive-diff · percentify · itables · coola · narwhals · chispa · pyreadstat · qpd