$npx skillfedfor your agent

boilerpy3

Python port of Boilerpipe, for HTML boilerplate removal and text extraction

With conditionsPyPI UtilitiesReleased Nov 2023207.8K downloads / moApache 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — boilerpy3-1.0.7-py3-none-any.whl
v1.0.7 · released 2023-11-01 · Python >=3.6

Yes, if you need straightforward HTML boilerplate removal and your use case tolerates a dormant library. The package is stable, dependency-free, and widely used, making it a low-risk choice for text extraction. However, do not install if you require active maintenance, support for modern web frameworks, or performance guarantees on contemporary HTML patterns—consider alternatives if those are critical.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with no runtime dependencies.
  • Dormant maintenance status (last commit 2024-08-20, no release since 2023-11-01) means bug fixes and updates are unlikely, though the package remains functional for its core use case.

License · maintenance · safety

Apache 2.0 (permissive) — Apache 2.0 permissive license allows free use, modification, and distribution with minimal restrictions—suitable for most commercial and open-source projects.

last release 2023-11-01 (1017 days) · last repo commit 2024-08-20 · 96 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 207,848 downloads/mo, #9,541 on PyPI

Verify before relying

pip install boilerpy3

from boilerpy3 import extractors

extractor = extractors.ArticleExtractor()
content = extractor.get_content('<html>...</html>')
  • How well the Boilerpipe 1.2-equivalent algorithm performs on modern web layouts compared to more recent extraction libraries.
  • Whether the package handles edge cases in malformed or unusual HTML gracefully when raise_on_failure is False.
  • Current download volume and active user base given the dormant maintenance status.
Same gist for agents: .md · .json

What it is and what it does

BoilerPy3 is a Python port of the Boilerpipe library, a text extraction tool that identifies and extracts the main content from HTML pages while discarding boilerplate elements like navigation, advertisements, and sidebars. It provides multiple extractors tuned for different scenarios—ArticleExtractor for news articles, DefaultExtractor for generic pages, and others for specialized use cases. The library works by analyzing HTML structure and text density to classify blocks as content or boilerplate.

The package has no runtime dependencies and supports Python 3.6 and later. It offers three main interfaces: extracting plain text, extracting marked HTML chunks, or retrieving a document object with additional metadata like title. The documentation notes that while URL-fetching methods are provided, using an external HTTP library like Requests is recommended for production use.

Use it for

  • Extract article text from news websites or blog posts for content aggregation or archival systems.
  • Clean HTML before feeding it to NLP pipelines or text analysis tools that need main content only.
  • Build web scrapers that need to isolate article body from page chrome and advertisements.
  • Preprocess HTML for machine learning training data where boilerplate noise would degrade model quality.
  • Implement fallback content extraction when more specialized algorithms fail (via raise_on_failure parameter).

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need straightforward HTML boilerplate removal and your use case tolerates a dormant library.

The package is stable, dependency-free, and widely used, making it a low-risk choice for text extraction. However, do not install if you require active maintenance, support for modern web frameworks, or performance guarantees on contemporary HTML patterns—consider alternatives if those are critical.

Install

boilerpy3 on PyPI

Before you install

Low install friction with no runtime dependencies. Dormant maintenance status (last commit 2024-08-20, no release since 2023-11-01) means bug fixes and updates are unlikely, though the package remains functional for its core use case.

License in practice

Apache 2.0 permissive license allows free use, modification, and distribution with minimal restrictions—suitable for most commercial and open-source projects.

Quickstart

pip install boilerpy3

from boilerpy3 import extractors

extractor = extractors.ArticleExtractor()
content = extractor.get_content('<html>...</html>')

Verify before relying

  • How well the Boilerpipe 1.2-equivalent algorithm performs on modern web layouts compared to more recent extraction libraries.
  • Whether the package handles edge cases in malformed or unusual HTML gracefully when raise_on_failure is False.
  • Current download volume and active user base given the dormant maintenance status.

Package facts

LicenseApache 2.0 permissive
Python supportSupports the current Python release >=3.6
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceDormant 1,017 days since the last release
Last repo commit
First released
Downloads207,848 / month, #9,541 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Utilities

Evidence: boilerpy3-1.0.7-py3-none-any.whl

Tags

Capabilities
html text extractionboilerplate removalarticle content extractionweb page text extractionnews article extractionhtml cleaningmain content extraction
Topics
content-extractionhtml-parsingweb-scraping
PyPI keywords
boilerpipeboilerpyhtml text extractiontext extractionfull text extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “html text extraction”

  • boilerpy3Extracts main article text and content from HTML, filtering out…
  • html-textExtracts plain text from HTML while filtering out styles, scripts,…
  • inscriptisConverts HTML documents to plain text while preserving layout,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Utilities packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
Pygments Worth it
PyPI · Utilities · released Mar 2026

Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.

Install it if you need to display or transform source code.

BSD-2-Clausepure Python · 3.9+
1.3Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo

See also goose3 · readable-content · UnityPy · newspaper4k · newspaper3k · html-text · trafilatura · date-guesser · extractcode · gnews