boilerpy3
Python port of Boilerpipe, for HTML boilerplate removal and text extraction
Decision gist · record as of 2026-08-14
Yes, if you need straightforward HTML boilerplate removal and your use case tolerates a dormant library. The package is stable, dependency-free, and widely used, making it a low-risk choice for text extraction. However, do not install if you require active maintenance, support for modern web frameworks, or performance guarantees on contemporary HTML patterns—consider alternatives if those are critical.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with no runtime dependencies.
- Dormant maintenance status (last commit 2024-08-20, no release since 2023-11-01) means bug fixes and updates are unlikely, though the package remains functional for its core use case.
License · maintenance · safety
Apache 2.0 (permissive) — Apache 2.0 permissive license allows free use, modification, and distribution with minimal restrictions—suitable for most commercial and open-source projects.
last release 2023-11-01 (1017 days) · last repo commit 2024-08-20 · 96 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 207,848 downloads/mo, #9,541 on PyPI
Alternatives
Verify before relying
pip install boilerpy3
from boilerpy3 import extractors
extractor = extractors.ArticleExtractor()
content = extractor.get_content('<html>...</html>')- How well the Boilerpipe 1.2-equivalent algorithm performs on modern web layouts compared to more recent extraction libraries.
- Whether the package handles edge cases in malformed or unusual HTML gracefully when raise_on_failure is False.
- Current download volume and active user base given the dormant maintenance status.
What it is and what it does
BoilerPy3 is a Python port of the Boilerpipe library, a text extraction tool that identifies and extracts the main content from HTML pages while discarding boilerplate elements like navigation, advertisements, and sidebars. It provides multiple extractors tuned for different scenarios—ArticleExtractor for news articles, DefaultExtractor for generic pages, and others for specialized use cases. The library works by analyzing HTML structure and text density to classify blocks as content or boilerplate.
The package has no runtime dependencies and supports Python 3.6 and later. It offers three main interfaces: extracting plain text, extracting marked HTML chunks, or retrieving a document object with additional metadata like title. The documentation notes that while URL-fetching methods are provided, using an external HTTP library like Requests is recommended for production use.
Use it for
- Extract article text from news websites or blog posts for content aggregation or archival systems.
- Clean HTML before feeding it to NLP pipelines or text analysis tools that need main content only.
- Build web scrapers that need to isolate article body from page chrome and advertisements.
- Preprocess HTML for machine learning training data where boilerplate noise would degrade model quality.
- Implement fallback content extraction when more specialized algorithms fail (via raise_on_failure parameter).
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need straightforward HTML boilerplate removal and your use case tolerates a dormant library.
The package is stable, dependency-free, and widely used, making it a low-risk choice for text extraction. However, do not install if you require active maintenance, support for modern web frameworks, or performance guarantees on contemporary HTML patterns—consider alternatives if those are critical.
Install
boilerpy3 on PyPI
Before you install
Low install friction with no runtime dependencies. Dormant maintenance status (last commit 2024-08-20, no release since 2023-11-01) means bug fixes and updates are unlikely, though the package remains functional for its core use case.
License in practice
Apache 2.0 permissive license allows free use, modification, and distribution with minimal restrictions—suitable for most commercial and open-source projects.
Quickstart
pip install boilerpy3
from boilerpy3 import extractors
extractor = extractors.ArticleExtractor()
content = extractor.get_content('<html>...</html>')
Verify before relying
- How well the Boilerpipe 1.2-equivalent algorithm performs on modern web layouts compared to more recent extraction libraries.
- Whether the package handles edge cases in malformed or unusual HTML gracefully when raise_on_failure is False.
- Current download volume and active user base given the dormant maintenance status.
Package facts
| License | Apache 2.0 permissive |
| Python support | Supports the current Python release >=3.6 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Dormant 1,017 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 207,848 / month, #9,541 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Utilities |
Evidence: boilerpy3-1.0.7-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html text extraction”
- boilerpy3Extracts main article text and content from HTML, filtering out…
- html-textExtracts plain text from HTML while filtering out styles, scripts,…
- inscriptisConverts HTML documents to plain text while preserving layout,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Utilities packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
See also goose3 · readable-content · UnityPy · newspaper4k · newspaper3k · html-text · trafilatura · date-guesser · extractcode · gnews