$npx skillfedfor your agent

tdda

Test-driven data analysis: command-line tools and Python APIs for data validation, testing analytical pipelines, automatic test generation and more.

With conditionsPyPI TestingReleased Jul 2026293.2K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — tdda-3.3.0-py3-none-any.whl
v3.3.0 · released 2026-07-13 · Python >=3.8 · 12 runtime deps: numpy, pandas, pyarrow, pyyaml, pytest, chardet, rich, regex

Yes, if you work with data pipelines and want automated validation. tdda fills a real gap between unit testing and data quality tools—reference testing catches regressions cheaply, and constraint discovery is faster than manual validation rules. Active maintenance, permissive license, and no security vulnerabilities. The 12-package dependency footprint is substantial but standard for data work; install it in projects where pandas or polars are already present.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.8; optional database support (PostgreSQL, MySQL/MariaDB, MongoDB) requires separate driver installation.
  • Low friction: pure Python wheel, active maintenance (last commit 2026-07-13), and a substantial dependency stack (numpy, pandas, pyarrow, pyyaml, pytest, chardet, rich, regex, tomli_w, tomli, polars, requests) that will be installed together.
  • Suitable for projects already using data science libraries.

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you can use, modify, and distribute tdda freely in commercial and private projects with minimal restrictions.

last release 2026-07-13 (32 days) · last repo commit 2026-07-13 · 310 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 293,247 downloads/mo, #7,955 on PyPI

Verify before relying

pip install tdda

from tdda.referencetest import ReferenceTestCase
import unittest

class MyDataTest(ReferenceTestCase):
    def test_pipeline(self):
        result = my_analysis_function()
        self.assertDataFrameCorrect(result, 'expected.csv')
  • Whether constraint discovery works on databases other than those listed in optional setup.
  • Performance characteristics when working with very large Parquet files or DataFrames.
  • Compatibility of CSVW and Frictionless metadata conversion with all format variants.
Same gist for agents: .md · .json

What it is and what it does

tdda is a Python framework for test-driven data analysis—a methodology where you write tests for data pipelines before or alongside the analysis code itself. It extends unittest and pytest with reference testing (comparing outputs to stored baselines), automatic test generation from command-line scripts, and constraint-based validation. The package also includes utilities for discovering data constraints from existing DataFrames or files, inferring regular expressions from string columns, diffing data across formats, and documenting CSV schemas in portable metadata files.

The core use case is validating analytical pipelines: you define what "correct" data looks like (via constraints or reference outputs), then verify that new data conforms to those rules. It integrates with pandas, polars, and parquet, and can work with relational databases when optional drivers are installed. The reference testing mode is particularly useful for regression detection—if your pipeline's output changes unexpectedly, the test catches it immediately.

Use it for

  • Write reference tests for data transformation pipelines to catch regressions when code or data changes.
  • Automatically generate baseline tests from existing command-line scripts or programs without manual test writing.
  • Discover and enforce data quality constraints (e.g., non-null columns, value ranges) on incoming datasets.
  • Infer regex patterns from sample string data to validate or extract structured text fields.
  • Compare two versions of a dataset (Parquet or CSV) and report row-level and column-level differences visually.
  • Document CSV and flat-file formats in portable metadata files compatible with CSVW and Frictionless standards.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you work with data pipelines and want automated validation.

tdda fills a real gap between unit testing and data quality tools—reference testing catches regressions cheaply, and constraint discovery is faster than manual validation rules. Active maintenance, permissive license, and no security vulnerabilities. The 12-package dependency footprint is substantial but standard for data work; install it in projects where pandas or polars are already present.

Install

tdda on PyPI

Before you install

Low friction: pure Python wheel, active maintenance (last commit 2026-07-13), and a substantial dependency stack (numpy, pandas, pyarrow, pyyaml, pytest, chardet, rich, regex, tomli_w, tomli, polars, requests) that will be installed together. Suitable for projects already using data science libraries.

Requires Python >=3.8; optional database support (PostgreSQL, MySQL/MariaDB, MongoDB) requires separate driver installation.

License in practice

MIT license is permissive; you can use, modify, and distribute tdda freely in commercial and private projects with minimal restrictions.

Quickstart

pip install tdda

from tdda.referencetest import ReferenceTestCase
import unittest

class MyDataTest(ReferenceTestCase):
    def test_pipeline(self):
        result = my_analysis_function()
        self.assertDataFrameCorrect(result, 'expected.csv')

Verify before relying

  • Whether constraint discovery works on databases other than those listed in optional setup.
  • Performance characteristics when working with very large Parquet files or DataFrames.
  • Compatibility of CSVW and Frictionless metadata conversion with all format variants.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
12 packages
numpypandaspyarrowpyyamlpytestchardetrichregextomli_wtomlipolarsrequests
MaintenanceActively maintained 32 days since the last release
Last repo commit
First released
Downloads293,247 / month, #7,955 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: tdda-3.3.0-py3-none-any.whl

Tags

Capabilities
data validation testingconstraint discovery pandasreference testing data pipelinesautomatic test generationdata quality constraintsregex inference from datadata diff comparison
Topics
data-validationtest-generationconstraint-discovery
PyPI keywords
tddaconstraintreferencetestrexpy

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “constraint discovery pandas”

  • tddatdda provides test-driven data analysis tools: reference testing for…
  • amplpyamplpy is a Python interface to AMPL, an algebraic modeling language…
  • docplexModeling library for building and solving mathematical optimization…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also pydeequ · pandas-schema · dtale · duckdb · pyddq · collate-data-diff · tdd-guard-pytest · cucumber-expressions · iregexp-check · pytest-html