psmpy
Propensity score matching for python and graphical plots
Decision gist · record as of 2026-08-14
Yes, if you are conducting observational studies in epidemiology or health research and need propensity score matching. The package is straightforward to use, well-integrated with the scientific Python stack, and permissively licensed. However, maintenance is aging (last release 275 days ago), so consider it stable for established workflows but not actively developed; for active support or cutting-edge features, evaluate alternatives or contribute to the repository.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction: pure Python wheel with six standard scientific dependencies (matplotlib, numpy, pandas, seaborn, scikit-learn, scipy).
- Maintenance status is aging—last release 275 days ago, 62 repository stars, no recent activity—so expect slower response to issues.
License · maintenance · safety
permissive license (permissive) — MIT license (permissive) means you can use, modify, and distribute PsmPy freely in commercial and private projects with minimal restrictions.
last release 2025-11-12 (275 days) · last repo commit 2025-11-12 · 62 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,663,512 downloads/mo, #2,952 on PyPI
Alternatives
Verify before relying
pip install psmpy
from psmpy import PsmPy
import pandas as pd
data = pd.read_csv('your_data.csv')
psm = PsmPy(data, treatment='treatment', indx='pat_id')
psm.logistic_ps(balance=True)
psm.kdtree_matched(matcher='propensity_logit', replacement=False)
matched_df = psm.df_matched- Whether the package handles missing data or requires complete-case analysis
- Performance characteristics with large datasets (sample size limits or memory requirements)
- Whether Python version support is truly unspecified or has implicit constraints
What it is and what it does
PsmPy is a Python library for propensity score matching, a statistical technique used in observational studies to estimate causal effects by matching treated and control subjects on their propensity scores—the predicted probability of receiving treatment given observed covariates. It reduces confounding bias by creating comparable treatment and control groups, mimicking the balance achieved in randomized trials.
The package provides logistic regression to compute propensity scores, KNN-based matching algorithms supporting both 1:1 and 1:many matching with optional caliper constraints, and visualization tools to assess covariate balance before and after matching. It integrates with pandas, scikit-learn, and matplotlib, and includes effect size calculations (Cohen's D) to quantify the standardized mean differences between groups.
Use it for
- Estimate treatment effects in retrospective cohort studies where randomization is infeasible or unethical
- Assess whether an intervention (e.g., a medical procedure) causally affects outcomes (e.g., mortality, length of stay) in observational data
- Balance covariates between treated and control groups to reduce confounding in epidemiological analyses
- Visualize covariate imbalance before and after matching to validate the matching procedure
- Compare multiple matching strategies (1:1 vs. 1:many, with/without replacement) on the same dataset
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are conducting observational studies in epidemiology or health research and need propensity score matching.
The package is straightforward to use, well-integrated with the scientific Python stack, and permissively licensed. However, maintenance is aging (last release 275 days ago), so consider it stable for established workflows but not actively developed; for active support or cutting-edge features, evaluate alternatives or contribute to the repository.
Install
psmpy on PyPI
Before you install
Low friction: pure Python wheel with six standard scientific dependencies (matplotlib, numpy, pandas, seaborn, scikit-learn, scipy). Maintenance status is aging—last release 275 days ago, 62 repository stars, no recent activity—so expect slower response to issues.
License in practice
MIT license (permissive) means you can use, modify, and distribute PsmPy freely in commercial and private projects with minimal restrictions.
Quickstart
pip install psmpy
from psmpy import PsmPy
import pandas as pd
data = pd.read_csv('your_data.csv')
psm = PsmPy(data, treatment='treatment', indx='pat_id')
psm.logistic_ps(balance=True)
psm.kdtree_matched(matcher='propensity_logit', replacement=False)
matched_df = psm.df_matched
Verify before relying
- Whether the package handles missing data or requires complete-case analysis
- Performance characteristics with large datasets (sample size limits or memory requirements)
- Whether Python version support is truly unspecified or has implicit constraints
Package facts
| License | permissive license permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesmatplotlibnumpypandasseabornscikit-learnscipy |
| Maintenance | Aging 275 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,663,512 / month, #2,952 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: psmpy-0.3.16-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “propensity score matching python”
- psmpyPsmPy implements propensity score matching for observational studies,…
- cemImplements Coarsened Exact Matching (CEM), a statistical matching…
- pfzyProvides fuzzy string matching with match indices using the fzy…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also causallib · cem · empirical-calibration · econml · causalml · statsmodels · pyhdfe · pystan · spglm · dowhy