skrub
Machine learning with dataframes
Decision gist · record as of 2026-08-14
Yes. skrub is actively maintained, has no known vulnerabilities, low install friction, and fills a genuine gap in the sklearn ecosystem for dataframe-native preprocessing. It is well-suited for teams doing tabular machine learning with pandas and scikit-learn. Install it if you regularly work with raw dataframes and want to avoid writing custom preprocessing code.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; scikit-learn, pandas, and numpy must be installed.
- Low friction: pure Python wheel with well-established dependencies (numpy, pandas, scikit-learn, scipy).
- Active maintenance with a recent release 39 days ago and steady repository activity.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause is permissive; you can use, modify, and distribute skrub freely in commercial and private projects with minimal restrictions.
last release 2026-07-06 (39 days) · last repo commit 2026-08-12 · 1,647 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 153,337 downloads/mo, #10,884 on PyPI
Alternatives
Verify before relying
pip install skrub
import skrub
from skrub import Joiner
# Use skrub transformers in a sklearn Pipeline or standalone- Specific transformers and their capabilities (e.g., handling missing values, encoding strategies, joining logic) are not detailed in the fact sheet.
- Performance characteristics on large datasets or typical dataframe sizes are not documented here.
- Whether skrub integrates directly into sklearn Pipelines or requires wrapper code is not explicit in the fact sheet.
What it is and what it does
skrub is a Python library that bridges the gap between raw dataframes and machine learning models by providing transformers and utilities for common data preparation tasks. It sits in the scikit-learn ecosystem and works with pandas DataFrames and numpy arrays, offering tools for feature engineering, encoding, and data cleaning that are typically needed before training models.
The library depends on numpy, pandas, scikit-learn, scipy, and visualization tools (matplotlib, pydot) for its operations. It is actively maintained, supports Python 3.10 through 3.14, and has been in production use since late 2023. The package is designed to integrate with sklearn's Pipeline API and other standard ML workflows, making it a natural fit for teams already using those tools.
Use it for
- Prepare messy tabular data with missing values and mixed data types for supervised learning.
- Encode categorical features and handle string columns in a sklearn-compatible way.
- Join multiple dataframes and align them for feature engineering in a machine learning pipeline.
- Transform raw CSV or database exports into clean feature matrices ready for model training.
- Build reproducible data preprocessing workflows that integrate with sklearn Pipelines.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
skrub is actively maintained, has no known vulnerabilities, low install friction, and fills a genuine gap in the sklearn ecosystem for dataframe-native preprocessing. It is well-suited for teams doing tabular machine learning with pandas and scikit-learn. Install it if you regularly work with raw dataframes and want to avoid writing custom preprocessing code.
Install
skrub on PyPI
Before you install
Low friction: pure Python wheel with well-established dependencies (numpy, pandas, scikit-learn, scipy). Active maintenance with a recent release 39 days ago and steady repository activity.
Requires Python 3.10 or later; scikit-learn, pandas, and numpy must be installed.
License in practice
BSD-3-Clause is permissive; you can use, modify, and distribute skrub freely in commercial and private projects with minimal restrictions.
Quickstart
pip install skrub
import skrub
from skrub import Joiner
# Use skrub transformers in a sklearn Pipeline or standalone
Verify before relying
- Specific transformers and their capabilities (e.g., handling missing values, encoding strategies, joining logic) are not detailed in the fact sheet.
- Performance characteristics on large datasets or typical dataframe sizes are not documented here.
- Whether skrub integrates directly into sklearn Pipelines or requires wrapper code is not explicit in the fact sheet.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagesnumpypandasscikit-learnscipyjinja2matplotlibrequestspydot |
| Maintenance | Actively maintained 39 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 153,337 / month, #10,884 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/EngineeringTopic :: Software Development :: Libraries |
Evidence: skrub-0.10.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “dataframe preprocessing for machine learning”
- skrubskrub prepares and transforms dataframes for machine learning by…
- scikit-learnscikit-learn provides a comprehensive Python library for supervised…
- grizzgrizz provides composable ingestors and transformers to load and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also pyjanitor · sklearn-pandas · sagemaker-scikit-learn-extension · gspread-pandas · feature-engine · bigframes · sklearndf · scikit-learn · pandas · miceforest