cem
Coarsened Exact Matching for Causal Inference
Decision gist · record as of 2026-08-14
Yes, if you are conducting causal inference with observational data and need a lightweight, dependency-minimal matching technique. The package has no known vulnerabilities and low install friction. However, the dormant maintenance status and unclear license terms warrant verification before production use. Suitable for research and academic work; exercise caution in proprietary contexts until license is clarified.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.9,<3.13; pandas and numpy must be installed.
- Low friction installation with only pandas and numpy as runtime dependencies.
- Package is dormant (last release 2023-10-12) with no recent maintenance signal, though it remains compatible with Python 3.9–3.12.
License · maintenance · safety
(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.
last release 2023-10-12 (1037 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,653,150 downloads/mo, #2,958 on PyPI
Alternatives
Verify before relying
pip install cem
from cem.match import match
from cem.coarsen import coarsen
from cem.imbalance import L1
import pandas as pd
X_coarse = coarsen(X, T, "l1")
weights = match(X_coarse, T)
imbalance = L1(X_coarse, weights)- Whether the dormant status signals abandonment risk or stable maintenance.
- Actual license terms and conditions (SPDX identifier and full license text not provided).
- Performance characteristics and scalability limits for large datasets.
- How CEM results compare to other matching techniques in practice.
What it is and what it does
CEM is a lightweight Python library for Coarsened Exact Matching, a statistical method that improves causal inference by reducing covariate imbalance in observational data. It works by coarsening continuous and categorical predictor variables into strata, then matching or reweighting observations to balance treatment and control groups. The library provides automatic and manual coarsening workflows, implements L1 and L2 multivariate imbalance measures, and produces observation weights suitable for downstream regression analysis.
The package is designed for researchers and analysts working with observational studies who need treatment effect estimates that are robust to model specification. It depends only on pandas and numpy, making it lightweight and easy to integrate into existing data analysis pipelines. Users typically coarsen their data, apply matching to generate weights, measure resulting imbalance, and use those weights in weighted regression models.
Use it for
- Reduce covariate imbalance in observational studies before estimating treatment effects.
- Generate observation weights for weighted regression to improve causal inference robustness.
- Compare L1 and L2 imbalance measures across different coarsening strategies.
- Preprocess data for causal inference when alternative matching is not suitable.
- Validate that matched samples have acceptable covariate balance before analysis.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are conducting causal inference with observational data and need a lightweight, dependency-minimal matching technique.
The package has no known vulnerabilities and low install friction. However, the dormant maintenance status and unclear license terms warrant verification before production use. Suitable for research and academic work; exercise caution in proprietary contexts until license is clarified.
Install
cem on PyPI
Before you install
Low friction installation with only pandas and numpy as runtime dependencies. Package is dormant (last release 2023-10-12) with no recent maintenance signal, though it remains compatible with Python 3.9–3.12.
Requires Python >=3.9,<3.13; pandas and numpy must be installed.
License in practice
License treatment is unclear—no SPDX identifier or raw license text is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.
Quickstart
pip install cem
from cem.match import match
from cem.coarsen import coarsen
from cem.imbalance import L1
import pandas as pd
X_coarse = coarsen(X, T, "l1")
weights = match(X_coarse, T)
imbalance = L1(X_coarse, weights)
Verify before relying
- Whether the dormant status signals abandonment risk or stable maintenance.
- Actual license terms and conditions (SPDX identifier and full license text not provided).
- Performance characteristics and scalability limits for large datasets.
- How CEM results compare to other matching techniques in practice.
Package facts
| License | Not declared unclear |
| Python support | Capped below the current Python release >=3.9,<3.13 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespandasnumpy |
| Maintenance | Dormant 1,037 days since the last release |
| First released | |
| Downloads | 2,653,150 / month, #2,958 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.9 |
Evidence: cem-1.1.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “coarsened exact matching”
- cemImplements Coarsened Exact Matching (CEM), a statistical matching…
- momentchi2Computes the cumulative distribution function of a weighted sum of…
- py-tlshGenerates locality-sensitive hashes for fuzzy matching of binary…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also psmpy · causalml · econml · dowhy · causallib · empirical-calibration · pyhdfe · imbalanced-learn · spreg · statsmodels