$npx skillfedfor your agent

readabilipy

Python wrapper for Mozilla's Readability.js

With conditionsPyPI HTMLReleased Dec 20241.5M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — readabilipy-0.3.0-py3-none-any.whl
v0.3.0 · released 2024-12-02 · Python >=3.6.0 · 4 runtime deps: beautifulsoup4, html5lib, lxml, regex

Yes, if you need reliable article extraction from HTML. The package is stable with low install friction and no known vulnerabilities. Dormant maintenance is typical for extraction tools that reach a stable state. Choose it for production use if you can accept either Node.js as a dependency or the pure-Python fallback; avoid it only if you need active feature development.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • To use Readability.js mode, Node.js version 14 or higher must be installed before installing readabilipy.
  • Python-only mode requires no external dependencies beyond the package's runtime dependencies.
  • Low install friction with a pure-Python fallback path.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice.

last release 2024-12-02 (620 days) · last repo commit 2024-12-02 · 359 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,490,968 downloads/mo, #3,843 on PyPI

Verify before relying

pip install readabilipy

from readabilipy import simple_json_from_html_string
article = simple_json_from_html_string(html_string, use_readability=False)
print(article['title'], article['plain_text'])
  • Whether Readability.js integration remains compatible with current Node.js versions given dormant maintenance status.
  • Performance and accuracy differences between Readability.js and pure-Python extraction modes on modern web content.
  • Real-world extraction quality on contemporary HTML structures and article formats.
Same gist for agents: .md · .json

What it is and what it does

ReadabiliPy wraps Mozilla's Readability.js Node.js package and provides a pure-Python alternative for extracting article content from HTML. It parses web pages and returns structured data including the article title, byline, simplified HTML, and plain text paragraphs. The package offers two extraction modes: one using Readability.js (requires Node.js 14 or higher) and one using a built-in Python parser that requires no external dependencies beyond beautifulsoup4, html5lib, lxml, and regex.

You can use it as a command-line tool to batch-process HTML files into JSON, or import it as a library in Python code. The library normalizes all text output using NFKC Unicode normalization and optionally adds SHA-256 content digests or hierarchical node indexes to the output structure.

Use it for

  • Extract article text from web pages for content analysis or archival without manual HTML parsing.
  • Build a scraper that converts HTML articles into structured JSON with title, author, and plain text paragraphs.
  • Process downloaded HTML files in batch via the command-line tool to generate article metadata and content.
  • Create a content pipeline that needs both structured HTML and plain text representations of articles.
  • Compare extraction quality between Readability.js and Python-only parser outputs on the same content.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need reliable article extraction from HTML.

The package is stable with low install friction and no known vulnerabilities. Dormant maintenance is typical for extraction tools that reach a stable state. Choose it for production use if you can accept either Node.js as a dependency or the pure-Python fallback; avoid it only if you need active feature development.

Install

readabilipy on PyPI

Before you install

Low install friction with a pure-Python fallback path. Dormant maintenance (last release 620 days ago, last commit 2024-12-02) but repository remains active. Supports Python 3.6 through 3.12.

To use Readability.js mode, Node.js version 14 or higher must be installed before installing readabilipy. Python-only mode requires no external dependencies beyond the package's runtime dependencies.

License in practice

MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license and copyright notice.

Quickstart

pip install readabilipy

from readabilipy import simple_json_from_html_string
article = simple_json_from_html_string(html_string, use_readability=False)
print(article['title'], article['plain_text'])

Verify before relying

  • Whether Readability.js integration remains compatible with current Node.js versions given dormant maintenance status.
  • Performance and accuracy differences between Readability.js and pure-Python extraction modes on modern web content.
  • Real-world extraction quality on contemporary HTML structures and article formats.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.6.0
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
beautifulsoup4html5liblxmlregex
MaintenanceDormant 620 days since the last release
Last repo commit
First released
Downloads1,490,968 / month, #3,843 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy

Evidence: readabilipy-0.3.0-py3-none-any.whl

Tags

Capabilities
extract article from htmlreadability parser pythonweb content extractionhtml to plain textarticle scraping librarymozilla readability wrappercontent extraction tool
Topics
web-scrapingcontent-extractionhtml-parsing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “html to plain text”

  • readabilipyExtracts article content from HTML using either Mozilla's…
  • html-textExtracts plain text from HTML while filtering out styles, scripts,…
  • html2textConverts HTML to clean, readable plain text or Markdown-formatted…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also readability-lxml · breadability · readable-content · goose3 · sumy · newspaper3k · trafilatura · strip-markdown · htmldate