category-encoders
A package for encoding categorical variables for machine learning
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has low installation friction, carries no known vulnerabilities, and provides a well-documented suite of encoding methods that integrate directly into sklearn workflows. It solves a common data preprocessing problem with multiple proven techniques. Install it if you work with categorical data in machine learning projects.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction installation with a pure Python wheel.
- Depends on established numerical and statistical libraries (numpy, pandas, scikit-learn, scipy, statsmodels).
- Actively maintained with a recent release.
License · maintenance · safety
BSD-3 (permissive) — BSD-3 permissive license allows commercial and private use with minimal restrictions; attribution required.
last release 2026-07-26 (19 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 3,123,199 downloads/mo, #2,741 on PyPI
Alternatives
Verify before relying
pip install category_encoders
from category_encoders import BinaryEncoder
import pandas as pd
X = pd.DataFrame({'gender': ['male', 'female'], 'age': [25, 32]})
enc = BinaryEncoder(cols=['gender']).fit(X)
numeric_X = enc.transform(X)- Whether supervised encoders (TargetEncoder, LeaveOneOut) handle class imbalance or missing values automatically
- Performance characteristics when encoding high-cardinality categorical features with many unique values
- Memory overhead of different encoding methods on large datasets
What it is and what it does
Category-encoders is a scikit-learn-compatible library that converts categorical variables into numeric form using a range of encoding strategies. It provides both unsupervised methods (binary, one-hot, ordinal, hashing) that treat all categories equally, and supervised methods (target encoding, CatBoost encoding, leave-one-out) that leverage target information to create more predictive numeric representations. The library integrates seamlessly into sklearn pipelines and accepts numpy arrays or pandas dataframes as input.
The package is designed for machine learning workflows where categorical features must be converted to numeric form. It handles the common problem of deciding which encoding method to use by offering multiple techniques with different statistical properties and computational trade-offs. Supervised encoders can reduce overfitting through nested cross-validation wrappers, and the library supports both classification and regression tasks through specialized wrappers.
Use it for
- Encode categorical features like gender or country in a classification pipeline using one-hot or binary encoding
- Apply target encoding to high-cardinality categorical variables to capture predictive signal from the target variable
- Use LeaveOneOut encoding in supervised settings to prevent overfitting while maintaining target information
- Build feature engineering pipelines that automatically encode all object-type columns in a pandas dataframe
- Compare multiple encoding strategies on the same dataset to select the best performing method
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has low installation friction, carries no known vulnerabilities, and provides a well-documented suite of encoding methods that integrate directly into sklearn workflows. It solves a common data preprocessing problem with multiple proven techniques. Install it if you work with categorical data in machine learning projects.
Install
category-encoders on PyPI
Before you install
Low friction installation with a pure Python wheel. Depends on established numerical and statistical libraries (numpy, pandas, scikit-learn, scipy, statsmodels). Actively maintained with a recent release.
License in practice
BSD-3 permissive license allows commercial and private use with minimal restrictions; attribution required.
Quickstart
pip install category_encoders
from category_encoders import BinaryEncoder
import pandas as pd
X = pd.DataFrame({'gender': ['male', 'female'], 'age': [25, 32]})
enc = BinaryEncoder(cols=['gender']).fit(X)
numeric_X = enc.transform(X)
Verify before relying
- Whether supervised encoders (TargetEncoder, LeaveOneOut) handle class imbalance or missing values automatically
- Performance characteristics when encoding high-cardinality categorical features with many unique values
- Memory overhead of different encoding methods on large datasets
Package facts
| License | BSD-3 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesnumpypandaspatsyscikit-learnscipystatsmodels |
| Maintenance | Actively maintained 19 days since the last release |
| First released | |
| Downloads | 3,123,199 / month, #2,741 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: Other/Proprietary LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: category_encoders-2.10.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “categorical variable encoding”
- category-encodersTransforms categorical variables into numeric representations using…
- feature-engineFeature-engine provides transformers for engineering, selecting, and…
- sagemaker-scikit-learn-extensionExtends scikit-learn with additional estimators and preprocessing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also catboost · formulaic-contrasts · phik · datasieve · kmodes · sagemaker-scikit-learn-extension · ngboost · sklearndf · graycode · gender-guesser