emrvalidator
A Data Validation Tool for Healthcare Data
Decision gist · record as of 2026-08-14
Yes, if you are validating healthcare or clinical datasets and want a lightweight, pandas-based alternative to heavier frameworks. The low install friction and healthcare-specific validators make it well-suited for ETL and data warehouse contexts. However, note the aging maintenance status (200 days since last release) and verify Python version support before committing to production use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires pandas; Python version support is unspecified in the package metadata, so verify compatibility with your environment before use.
- Low friction: pure Python wheel with only pandas as a runtime dependency.
- Maintenance status is aging—last release 200 days ago—so expect slower response to issues or feature requests.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package freely in commercial or private projects with minimal restrictions.
last release 2026-01-26 (200 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 242,555 downloads/mo, #8,852 on PyPI
Alternatives
Verify before relying
pip install emrvalidator
from emrvalidator import DataValidator
import pandas as pd
df = pd.read_csv('patient_data.csv')
validator = DataValidator("Patient Data Quality")
validator.load_data(df)
validator.expect_column_exists('mrn').expect_mrn_format('mrn')
if validator.is_valid():
print("✓ Validations passed")- Whether numpy is truly optional or a transitive dependency of pandas that users should be aware of
- Actual performance comparison methodology behind the claimed 5-7x speedup over Great Expectations
- Python version support (requires_python is unspecified in the fact sheet)
- Whether the package is actively maintained or in maintenance-only mode given the 200-day gap since last release
What it is and what it does
EMRValidator is a Python data validation library designed for healthcare and clinical datasets. It provides a fluent API for chaining validation rules, with built-in support for healthcare-specific formats like MRN and ICD codes, alongside standard checks for nulls, ranges, uniqueness, and date formats. The library depends only on pandas and includes data profiling and HTML/JSON report generation.
It positions itself as a lighter alternative to Great Expectations, targeting healthcare ETL pipelines, data warehouses, and clinical analytics workflows. The package offers multiple API styles—fluent chaining, expectation suites, and reusable rule sets—and includes pre-built rule sets for common healthcare scenarios like patient demographics and financial data validation.
Use it for
- Validate patient demographics, MRN formats, and clinical codes in ETL pipelines before loading to a data warehouse
- Generate data quality reports and profiling summaries for healthcare datasets to identify missing or malformed records
- Define and reuse custom validation rule sets for claims data, encounter records, or revenue cycle management workflows
- Run real-time data quality checks on incoming clinical data in Airflow, dbt, or LLM pipeline contexts
- Validate ICD-9 and ICD-10 diagnosis codes and other healthcare-specific formats in bulk data migrations
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are validating healthcare or clinical datasets and want a lightweight, pandas-based alternative to heavier frameworks.
The low install friction and healthcare-specific validators make it well-suited for ETL and data warehouse contexts. However, note the aging maintenance status (200 days since last release) and verify Python version support before committing to production use.
Install
emrvalidator on PyPI
Before you install
Low friction: pure Python wheel with only pandas as a runtime dependency. Maintenance status is aging—last release 200 days ago—so expect slower response to issues or feature requests.
Requires pandas; Python version support is unspecified in the package metadata, so verify compatibility with your environment before use.
License in practice
MIT license is permissive; you can use, modify, and distribute this package freely in commercial or private projects with minimal restrictions.
Quickstart
pip install emrvalidator
from emrvalidator import DataValidator
import pandas as pd
df = pd.read_csv('patient_data.csv')
validator = DataValidator("Patient Data Quality")
validator.load_data(df)
validator.expect_column_exists('mrn').expect_mrn_format('mrn')
if validator.is_valid():
print("✓ Validations passed")
Verify before relying
- Whether numpy is truly optional or a transitive dependency of pandas that users should be aware of
- Actual performance comparison methodology behind the claimed 5-7x speedup over Great Expectations
- Python version support (requires_python is unspecified in the fact sheet)
- Whether the package is actively maintained or in maintenance-only mode given the 200-day gap since last release
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagepandas |
| Maintenance | Aging 200 days since the last release |
| First released | |
| Downloads | 242,555 / month, #8,852 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: emrvalidator-1.0.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “healthcare data validation”
- emrvalidatorEMRValidator is a data validation library specialized for healthcare…
- fhir.resourcesProvides Python classes and validation for all FHIR resource types…
- fhir-coreProvides Pydantic V2-based abstract base classes and primitive…
Give your agent the search over MCP, or paste the wish link into any chat.
More Quality Assurance packages
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.
Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.
Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.
pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.
Install it if your test suite takes long enough that parallelization would save meaningful time.
Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.
Install it if you work with CloudFormation templates.
See also carelytics · validations-engine · simple-icd-10-cm · icd-mappings · openmed · spark-expectations · hl7 · fhir.resources · airflow-provider-great-expectations · bids-validator