$npx skillfedfor your agent

missingpy

Missing Data Imputation for Python

With conditionsPyPI Scientific/EngineeringReleased Dec 2018296.2K downloads / mocopyleft licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — missingpy-0.2.0-py3-none-any.whl
v0.2.0 · released 2018-12-10

Yes, if you need a lightweight, zero-dependency imputation tool and can tolerate dormant maintenance. The scikit-learn API is familiar and the code is straightforward. No, if you require active maintenance, compatibility with recent library versions, or timely bug fixes. Test thoroughly on your specific data and Python version before production use.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with no runtime dependencies.
  • Maintenance is dormant—last release was 2018-12-10 and last commit 2024-02-29, so expect no active development or timely bug fixes.

License · maintenance · safety

copyleft license (copyleft) — Licensed under GPLv3 (copyleft). Any derivative work or bundled distribution must also be open-source under compatible terms.

last release 2018-12-10 (2804 days) · last repo commit 2024-02-29 · 246 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 296,221 downloads/mo, #7,901 on PyPI

Verify before relying

pip install missingpy

from missingpy import KNNImputer

X = [[1, 2, nan], [3, 4, 3], [nan, 6, 5], [8, 8, 7]]
imputer = KNNImputer(n_neighbors=2)
X_imputed = imputer.fit_transform(X)
  • Whether the package works with current versions of common numerical libraries (last release predates many ecosystem updates)
  • Performance characteristics on large datasets or high-dimensional data
  • Whether categorical variable support in MissForest is fully documented and stable
Same gist for agents: .md · .json

What it is and what it does

missingpy provides two methods for imputing missing values in numerical data arrays: k-Nearest Neighbors (KNNImputer) and Random Forest (MissForest). Both follow scikit-learn's fit/transform interface, making them familiar to users of that ecosystem. KNNImputer replaces missing values by averaging values from the k nearest neighbors; MissForest uses iterative random forest predictions, starting with the column containing the fewest missing values and working outward. The library handles both numerical and categorical variables (in MissForest) and allows configuration of neighbor counts, distance metrics, and missing-value thresholds.

The package has no external runtime dependencies, making installation straightforward. However, it has been dormant since late 2018—the latest release is 0.2.0 from December 2018, and while the repository shows a commit in February 2024, there have been no new releases. This means the codebase may not be compatible with recent versions of common libraries, and bug reports or feature requests are unlikely to receive timely attention.

Use it for

  • Preprocess datasets with missing values before feeding them into machine learning pipelines that require complete data.
  • Impute missing entries in time-series or sensor data where nearest-neighbor patterns reflect realistic local structure.
  • Handle mixed numerical and categorical missing data in tabular datasets using MissForest's iterative approach.
  • Replace missing values in microarray or genomic data where the k-NN methodology has established validity.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need a lightweight, zero-dependency imputation tool and can tolerate dormant maintenance.

The scikit-learn API is familiar and the code is straightforward. No, if you require active maintenance, compatibility with recent library versions, or timely bug fixes. Test thoroughly on your specific data and Python version before production use.

Install

missingpy on PyPI

Before you install

Low install friction with no runtime dependencies. Maintenance is dormant—last release was 2018-12-10 and last commit 2024-02-29, so expect no active development or timely bug fixes.

License in practice

Licensed under GPLv3 (copyleft). Any derivative work or bundled distribution must also be open-source under compatible terms.

Quickstart

pip install missingpy

from missingpy import KNNImputer

X = [[1, 2, nan], [3, 4, 3], [nan, 6, 5], [8, 8, 7]]
imputer = KNNImputer(n_neighbors=2)
X_imputed = imputer.fit_transform(X)

Verify before relying

  • Whether the package works with current versions of common numerical libraries (last release predates many ecosystem updates)
  • Performance characteristics on large datasets or high-dimensional data
  • Whether categorical variable support in MissForest is fully documented and stable

Package facts

Licensecopyleft license copyleft
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceDormant 2,804 days since the last release
Last repo commit
First released
Downloads296,221 / month, #7,901 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: GNU General Public License v3 (GPLv3)Operating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: missingpy-0.2.0-py3-none-any.whl

Tags

Capabilities
missing data imputationknn imputationrandom forest imputationhandle missing valuesdata preprocessing missingmissforestfill nan values
Topics
data-preprocessingimputationscikit-learn-compatible

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “missing data imputation”

  • missingpyFills missing values in data arrays using k-Nearest Neighbors or…
  • miceforestPerforms Multiple Imputation by Chained Equations (MICE) using…
  • time-aware-imputerFills missing values in time-series data while respecting irregular…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also miceforest · time-aware-imputer · pynndescent · sagemaker-scikit-learn-extension · forestci · quantile-forest · treeinterpreter · ydf · pypots · ai4ts