miceforest
Multiple Imputation by Chained Equations with LightGBM
Decision gist · record as of 2026-08-14
Yes, if you need production-grade MICE imputation with speed and flexibility. The low install friction, MIT license, and lack of known vulnerabilities make it a safe choice. The aging maintenance status (291 days since last release) is a minor concern but not a blocker—the package is stable and the repository remains active. Install if your workflow requires multiple imputation or if you're working with missing data in pandas/numpy and want LightGBM's speed over traditional methods.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; LightGBM, NumPy, Pandas, PyArrow, and SciPy must be installed.
- Low friction install with a pure-Python wheel.
- Maintenance is aging—last release was 291 days ago—but the repository remains active with 411 stars and no archived status.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely provided you include the license notice.
last release 2025-10-27 (291 days) · last repo commit 2025-10-27 · 411 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 90,756 downloads/mo, #13,564 on PyPI
Alternatives
Verify before relying
import miceforest as mf
import pandas as pd
# Create kernel with missing data
kernel = mf.ImputationKernel(data_with_missing_values, random_state=1)
# Run MICE algorithm
kernel.mice(2)
# Get completed dataset
imputed_data = kernel.complete_data()- Whether GPU training support is functional and what GPU libraries are required.
- Performance benchmarks comparing miceforest to other MICE implementations.
- Whether the package is actively maintained or in maintenance-only mode given the 291-day release gap.
What it is and what it does
miceforest implements Multiple Imputation by Chained Equations (MICE), a statistical method for handling missing data by creating multiple plausible imputed datasets. It uses LightGBM as the underlying predictive model, which provides speed and memory efficiency compared to traditional MICE implementations. The package handles both numeric and categorical data automatically and supports mean matching to preserve the distribution of imputed values.
The package is designed for both research and production use. You can create a single imputed dataset for quick analysis, or generate multiple imputed datasets to quantify uncertainty from missing values. It integrates with pandas and numpy, fits into scikit-learn pipelines, and allows you to train models on complete data and apply them to new datasets with missing values. Data can be imputed in place to reduce memory overhead, and trained kernels can be saved and reloaded for consistent imputation of new data.
Use it for
- Fill missing values in survey or medical datasets while preserving statistical properties for downstream analysis.
- Generate multiple imputed datasets to assess how missing-data uncertainty affects model predictions or statistical inference.
- Impute new, unseen data using models trained on a reference dataset without retraining.
- Preprocess data with missing values as part of a scikit-learn machine learning pipeline.
- Handle datasets with mixed numeric and categorical columns automatically without manual encoding.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need production-grade MICE imputation with speed and flexibility.
The low install friction, MIT license, and lack of known vulnerabilities make it a safe choice. The aging maintenance status (291 days since last release) is a minor concern but not a blocker—the package is stable and the repository remains active. Install if your workflow requires multiple imputation or if you're working with missing data in pandas/numpy and want LightGBM's speed over traditional methods.
Install
miceforest on PyPI
Before you install
Low friction install with a pure-Python wheel. Maintenance is aging—last release was 291 days ago—but the repository remains active with 411 stars and no archived status.
Requires Python 3.10 or later; LightGBM, NumPy, Pandas, PyArrow, and SciPy must be installed.
License in practice
MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely provided you include the license notice.
Quickstart
import miceforest as mf
import pandas as pd
# Create kernel with missing data
kernel = mf.ImputationKernel(data_with_missing_values, random_state=1)
# Run MICE algorithm
kernel.mice(2)
# Get completed dataset
imputed_data = kernel.complete_data()
Verify before relying
- Whether GPU training support is functional and what GPU libraries are required.
- Performance benchmarks comparing miceforest to other MICE implementations.
- Whether the package is actively maintained or in maintenance-only mode given the 291-day release gap.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release <4.0,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packageslightgbmnumpypandaspyarrowscipy |
| Maintenance | Aging 291 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 90,756 / month, #13,564 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Natural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9 |
Evidence: miceforest-6.0.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “missing value imputation”
- miceforestPerforms Multiple Imputation by Chained Equations (MICE) using…
- time-aware-imputerFills missing values in time-series data while respecting irregular…
- missingpyFills missing values in data arrays using k-Nearest Neighbors or…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also missingpy · time-aware-imputer · awkward-pandas · sagemaker-scikit-learn-extension · lightgbm · skrub · sklearn-pandas · sparse · sklearndf · pypots