Boruta
Python Implementation of Boruta Feature Selection
Decision gist · record as of 2026-08-14
Yes, with conditions. Boruta is well-established (1625 GitHub stars, 258914 monthly downloads) and has no known vulnerabilities. Install it if you need all-relevant feature selection with a scikit-learn-compatible interface and can accept aging maintenance (last release 2024-08-13). Not suitable if you require active, frequent updates or the latest algorithmic refinements.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires numpy arrays as input (not pandas DataFrames directly); convert with .values if needed.
- Low friction: pure Python wheel with only three standard dependencies (numpy, scipy, scikit-learn).
- Last release 2024-08-13; repository is active but maintenance status is aging with 731 days since release.
License · maintenance · safety
BSD 3 clause (permissive) — BSD 3 clause is permissive; you can use, modify, and distribute Boruta with minimal restrictions in commercial or private projects.
last release 2024-08-13 (731 days) · last repo commit 2025-11-13 · 1,625 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 258,914 downloads/mo, #8,422 on PyPI
Alternatives
Verify before relying
pip install Boruta
from boruta import BorutaPy
feat_selector = BorutaPy(estimator, n_estimators='auto', verbose=2, random_state=1)
feat_selector.fit(X, y)
X_filtered = feat_selector.transform(X)- Whether the package is actively maintained going forward (aging status suggests infrequent updates).
- Performance characteristics on datasets with thousands of features.
- Current state of the original R package and how closely this implementation tracks it.
What it is and what it does
Boruta is a Python implementation of an all-relevant feature selection algorithm that identifies every feature carrying information useful for prediction, rather than finding a minimal optimal subset. It wraps ensemble methods from scikit-learn and uses a shadow-feature comparison approach with statistical testing to rank and select features. Unlike minimal-optimal methods that depend on classifier choice, Boruta aims to discover all contributing factors to your target variable, making it useful for exploratory analysis and understanding data relationships.
The package provides a scikit-learn-compatible interface (fit, transform, fit_transform) and includes refinements over the original R implementation: automatic estimator selection, feature ranking, percentile-based thresholds instead of strict maximums, and a two-step multiple-testing correction. It depends on numpy, scipy, and scikit-learn, making installation straightforward.
Use it for
- Exploratory data analysis: identify all features related to your target to understand which factors influence your phenomenon.
- Biological or medical data: use relaxed thresholds (perc parameter) when Bonferroni correction is too harsh for your domain.
- Feature engineering validation: rank candidate features to guide which ones to engineer further or combine.
- Preprocessing before minimal-optimal methods: run Boruta first to reduce noise, then apply stricter selectors.
- Interpretability: obtain feature rankings and support masks to explain model decisions to stakeholders.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Boruta is well-established (1625 GitHub stars, 258914 monthly downloads) and has no known vulnerabilities. Install it if you need all-relevant feature selection with a scikit-learn-compatible interface and can accept aging maintenance (last release 2024-08-13). Not suitable if you require active, frequent updates or the latest algorithmic refinements.
Install
boruta on PyPI
Before you install
Low friction: pure Python wheel with only three standard dependencies (numpy, scipy, scikit-learn). Last release 2024-08-13; repository is active but maintenance status is aging with 731 days since release.
Requires numpy arrays as input (not pandas DataFrames directly); convert with .values if needed.
License in practice
BSD 3 clause is permissive; you can use, modify, and distribute Boruta with minimal restrictions in commercial or private projects.
Quickstart
pip install Boruta
from boruta import BorutaPy
feat_selector = BorutaPy(estimator, n_estimators='auto', verbose=2, random_state=1)
feat_selector.fit(X, y)
X_filtered = feat_selector.transform(X)
Verify before relying
- Whether the package is actively maintained going forward (aging status suggests infrequent updates).
- Performance characteristics on datasets with thousands of features.
- Current state of the original R package and how closely this implementation tracks it.
Package facts
| License | BSD 3 clause permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesnumpyscikit-learnscipy |
| Maintenance | Aging 731 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 258,914 / month, #8,422 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: Boruta-0.4.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “all-relevant features”
- BorutaBoruta performs all-relevant feature selection by identifying all…
- mrmr-selectionImplements the mRMR (minimum Redundancy - Maximum Relevance) feature…
- haystack-experimentalProvides archived experimental features for Haystack LLM framework…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also mlxtend · mrmr-selection · powershap · treeinterpreter · forestci · feature-engine · scikit-learn · tensorflow-decision-forests · yellowbrick · sagemaker-scikit-learn-extension