readabilipy
Python wrapper for Mozilla's Readability.js
Decision gist · record as of 2026-08-14
Yes, if you need reliable article extraction from HTML. The package is stable with low install friction and no known vulnerabilities. Dormant maintenance is typical for extraction tools that reach a stable state. Choose it for production use if you can accept either Node.js as a dependency or the pure-Python fallback; avoid it only if you need active feature development.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- To use Readability.js mode, Node.js version 14 or higher must be installed before installing readabilipy.
- Python-only mode requires no external dependencies beyond the package's runtime dependencies.
- Low install friction with a pure-Python fallback path.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice.
last release 2024-12-02 (620 days) · last repo commit 2024-12-02 · 359 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,490,968 downloads/mo, #3,843 on PyPI
Alternatives
Verify before relying
pip install readabilipy
from readabilipy import simple_json_from_html_string
article = simple_json_from_html_string(html_string, use_readability=False)
print(article['title'], article['plain_text'])- Whether Readability.js integration remains compatible with current Node.js versions given dormant maintenance status.
- Performance and accuracy differences between Readability.js and pure-Python extraction modes on modern web content.
- Real-world extraction quality on contemporary HTML structures and article formats.
What it is and what it does
ReadabiliPy wraps Mozilla's Readability.js Node.js package and provides a pure-Python alternative for extracting article content from HTML. It parses web pages and returns structured data including the article title, byline, simplified HTML, and plain text paragraphs. The package offers two extraction modes: one using Readability.js (requires Node.js 14 or higher) and one using a built-in Python parser that requires no external dependencies beyond beautifulsoup4, html5lib, lxml, and regex.
You can use it as a command-line tool to batch-process HTML files into JSON, or import it as a library in Python code. The library normalizes all text output using NFKC Unicode normalization and optionally adds SHA-256 content digests or hierarchical node indexes to the output structure.
Use it for
- Extract article text from web pages for content analysis or archival without manual HTML parsing.
- Build a scraper that converts HTML articles into structured JSON with title, author, and plain text paragraphs.
- Process downloaded HTML files in batch via the command-line tool to generate article metadata and content.
- Create a content pipeline that needs both structured HTML and plain text representations of articles.
- Compare extraction quality between Readability.js and Python-only parser outputs on the same content.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need reliable article extraction from HTML.
The package is stable with low install friction and no known vulnerabilities. Dormant maintenance is typical for extraction tools that reach a stable state. Choose it for production use if you can accept either Node.js as a dependency or the pure-Python fallback; avoid it only if you need active feature development.
Install
readabilipy on PyPI
Before you install
Low install friction with a pure-Python fallback path. Dormant maintenance (last release 620 days ago, last commit 2024-12-02) but repository remains active. Supports Python 3.6 through 3.12.
To use Readability.js mode, Node.js version 14 or higher must be installed before installing readabilipy. Python-only mode requires no external dependencies beyond the package's runtime dependencies.
License in practice
MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice.
Quickstart
pip install readabilipy
from readabilipy import simple_json_from_html_string
article = simple_json_from_html_string(html_string, use_readability=False)
print(article['title'], article['plain_text'])
Verify before relying
- Whether Readability.js integration remains compatible with current Node.js versions given dormant maintenance status.
- Performance and accuracy differences between Readability.js and pure-Python extraction modes on modern web content.
- Real-world extraction quality on contemporary HTML structures and article formats.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.6.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesbeautifulsoup4html5liblxmlregex |
| Maintenance | Dormant 620 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,490,968 / month, #3,843 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy |
Evidence: readabilipy-0.3.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html to plain text”
- readabilipyExtracts article content from HTML using either Mozilla's…
- html-textExtracts plain text from HTML while filtering out styles, scripts,…
- html2textConverts HTML to clean, readable plain text or Markdown-formatted…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also readability-lxml · breadability · readable-content · goose3 · sumy · newspaper3k · trafilatura · strip-markdown · htmldate