feature-engine
Feature engineering and selection package with Scikit-learn's fit transform functionality
Decision gist · record as of 2026-08-14
Yes. Feature-engine is actively maintained, has no known vulnerabilities, installs with low friction, and provides a comprehensive suite of production-ready transformers that integrate seamlessly with scikit-learn workflows. It is well-suited for anyone building machine learning pipelines who wants to avoid writing repetitive feature engineering code.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later.
- Low install friction with a pure Python wheel.
- Active maintenance with a recent release 168 days ago and 2267 GitHub stars.
License · maintenance · safety
BSD 3 clause (permissive) — BSD 3-clause permissive license allows commercial and private use with minimal restrictions.
last release 2026-02-27 (168 days) · last repo commit 2026-07-31 · 2,267 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 331,267 downloads/mo, #7,518 on PyPI
Alternatives
Verify before relying
pip install feature_engine
import pandas as pd
from feature_engine.encoding import RareLabelEncoder
data = pd.DataFrame({'var_A': ['A']*10 + ['B']*10 + ['C']*2 + ['D']*1})
encoder = RareLabelEncoder(tol=0.10, n_categories=3)
encoded = encoder.fit_transform(data)- Whether all transformer classes are documented with parameter details and use-case guidance.
- Performance characteristics when applied to large datasets or high-dimensional feature spaces.
- Compatibility behavior when chaining transformers with custom preprocessing pipelines.
What it is and what it does
Feature-engine is a Python library that wraps common feature engineering and selection tasks into scikit-learn-compatible transformers. It provides methods for handling missing data, encoding categorical variables, discretising continuous features, capping or removing outliers, transforming and scaling variables, creating new features from existing ones, and selecting the most informative features for modeling. The library depends on numpy, pandas, scikit-learn, scipy, and statsmodels to perform its transformations.
You use it by instantiating a transformer (e.g., RareLabelEncoder, MeanImputer, DropCorrelatedFeatures), calling fit() on training data to learn parameters, and then calling transform() on new data to apply the same transformation. This design integrates naturally into scikit-learn pipelines and cross-validation workflows, making it straightforward to build reproducible feature engineering workflows without writing custom code for each task.
Use it for
- Encode rare categorical values into a single 'Rare' category to reduce cardinality before modeling.
- Impute missing values using mean, median, or arbitrary strategies learned from training data.
- Remove or cap outliers using statistical methods like Winsorization before training.
- Create datetime-derived features like day-of-week or cyclical encodings from timestamp columns.
- Select the most predictive features using correlation, information value, or model-based elimination.
- Transform skewed variables using log, Box-Cox, or Yeo-Johnson transformations for normality.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Feature-engine is actively maintained, has no known vulnerabilities, installs with low friction, and provides a comprehensive suite of production-ready transformers that integrate seamlessly with scikit-learn workflows. It is well-suited for anyone building machine learning pipelines who wants to avoid writing repetitive feature engineering code.
Install
feature-engine on PyPI
Before you install
Low install friction with a pure Python wheel. Active maintenance with a recent release 168 days ago and 2267 GitHub stars. Supports modern Python versions 3.9 through 3.14.
Requires Python 3.9 or later.
License in practice
BSD 3-clause permissive license allows commercial and private use with minimal restrictions.
Quickstart
pip install feature_engine
import pandas as pd
from feature_engine.encoding import RareLabelEncoder
data = pd.DataFrame({'var_A': ['A']*10 + ['B']*10 + ['C']*2 + ['D']*1})
encoder = RareLabelEncoder(tol=0.10, n_categories=3)
encoded = encoder.fit_transform(data)
Verify before relying
- Whether all transformer classes are documented with parameter details and use-case guidance.
- Performance characteristics when applied to large datasets or high-dimensional feature spaces.
- Compatibility behavior when chaining transformers with custom preprocessing pipelines.
Package facts
| License | BSD 3 clause permissive |
| Python support | Supports the current Python release >=3.9.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesnumpypandasscikit-learnscipystatsmodels |
| Maintenance | Actively maintained 168 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 331,267 / month, #7,518 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: BSD LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9 |
Evidence: feature_engine-1.9.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “feature engineering transformers”
- feature-engineFeature-engine provides transformers for engineering, selecting, and…
- sagemaker-scikit-learn-extensionExtends scikit-learn with additional estimators and preprocessing…
- sklearndfWraps scikit-learn estimators to return pandas DataFrames instead of…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also datasieve · sagemaker-scikit-learn-extension · skforecast · pycaret · skrub · Boruta · sktime · transformer-lens · sentence-transformers · validations-engine