$npx skillfedfor your agent

breadability

Port of Readability HTML parser in Python

Worth itPyPI WWW/HTTPReleased Aug 2026100.7K downloads / moBSDPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — breadability-0.1.21-py2.py3-none-any.whl
v0.1.21 · released 2026-08-12 · 3 runtime deps: docopt, chardet, lxml

Yes. The package is actively maintained, has no known vulnerabilities, installs with low friction, and solves a well-defined problem. It is suitable for production use if you need reliable HTML-to-readable-content extraction. The BSD license poses no restrictions. Install only if you can satisfy the lxml build dependency; otherwise, consider it a straightforward choice.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • lxml requires libxml2-dev and libxslt-dev system libraries to compile; install via apt-get or equivalent before pip install.
  • Low friction install with a pure-Python wheel.
  • Active maintenance—last release 2 days ago.

License · maintenance · safety

BSD (permissive) — BSD license (permissive). You can use, modify, and distribute this package freely in both open-source and commercial projects with minimal restrictions.

last release 2026-08-12 (2 days) · last repo commit 2026-08-12 · 205 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 100,652 downloads/mo, #12,978 on PyPI

Verify before relying

from breadability.readable import Article

html_text = "<html>...</html>"
url = "http://example.com/article"
document = Article(html_text, url=url)
print(document.readable)
  • Whether the package handles modern HTML5 and contemporary web layouts as well as the original Arc90 readability algorithm.
  • Performance characteristics on large documents or high-volume extraction workloads.
  • Accuracy of content extraction on contemporary news sites and blogging platforms.
Same gist for agents: .md · .json

What it is and what it does

Breadability is a Python port of the Arc90 Readability algorithm, designed to parse HTML pages and extract the main article content while filtering out navigation, sidebars, ads, and other boilerplate. It uses a scoring system to identify the most likely content container, then cleans and returns readable HTML or text.

The package provides both a command-line interface and a Python API. It depends on docopt for CLI argument parsing, chardet for character encoding detection, and lxml for HTML parsing. The library is maintained and actively used in production tools. It supports Python 2.6 through 3.6 and offers options for debugging parse decisions, opening results in a browser, or returning full documents instead of fragments.

Use it for

  • Extract article text from news websites for content aggregation or archival tools.
  • Clean HTML before feeding it to NLP or text analysis pipelines.
  • Build a web scraper that reliably isolates main content across diverse site layouts.
  • Debug why a parser chose certain nodes as content using verbose scoring output.
  • Convert web pages to readable text for offline reading or accessibility.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, has no known vulnerabilities, installs with low friction, and solves a well-defined problem. It is suitable for production use if you need reliable HTML-to-readable-content extraction. The BSD license poses no restrictions. Install only if you can satisfy the lxml build dependency; otherwise, consider it a straightforward choice.

Install

breadability on PyPI

Before you install

Low friction install with a pure-Python wheel. Active maintenance—last release 2 days ago. Depends on lxml, which requires C headers (libxml2-dev, libxslt-dev) at build time on some systems, but this is a one-time setup cost.

lxml requires libxml2-dev and libxslt-dev system libraries to compile; install via apt-get or equivalent before pip install.

License in practice

BSD license (permissive). You can use, modify, and distribute this package freely in both open-source and commercial projects with minimal restrictions.

Quickstart

from breadability.readable import Article

html_text = "<html>...</html>"
url = "http://example.com/article"
document = Article(html_text, url=url)
print(document.readable)

Verify before relying

  • Whether the package handles modern HTML5 and contemporary web layouts as well as the original Arc90 readability algorithm.
  • Performance characteristics on large documents or high-volume extraction workloads.
  • Accuracy of content extraction on contemporary news sites and blogging platforms.

Package facts

LicenseBSD permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
docoptchardetlxml
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads100,652 / month, #12,978 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 2.6Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.2Programming Language :: Python :: 3.3Programming Language :: Python :: 3.4Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: Implementation :: CPythonTopic :: Internet :: WWW/HTTPTopic :: Software Development :: Pre-processorsTopic :: Text Processing :: FiltersTopic :: Text Processing :: Markup :: HTML

Evidence: breadability-0.1.21-py2.py3-none-any.whl

Tags

Capabilities
html content extractionreadability parser pythonarticle text extractionweb page cleaningboilerplate removalhtml parsing readabilitymain content extraction
Topics
content-extractionhtml-parsingweb-scraping
PyPI keywords
bookiebreadabilitycontentHTMLparsingreadabilityreadable

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “readability parser python”

  • breadabilityExtracts the main readable content from HTML pages by identifying and…
  • readabilipyExtracts article content from HTML using either Mozilla's…
  • readability-lxmlExtracts and cleans the main article text and title from HTML…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also jusText · readability-lxml · readabilipy · goose3 · newspaper3k · readable-content · newspaper4k · textract · htmlmin · textacy