$npx skillfedfor your agent

jusText

Heuristic based boilerplate removal tool

With conditionsPyPI WWW/HTTPReleased Feb 202511.7M downloads / moThe BSD 2-Clause LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — justext-3.0.2-py2.py3-none-any.whl
v3.0.2 · released 2025-02-25 · 2 runtime deps: lxml, backports.functools-lru-cache

Yes, with conditions. The package is stable (Production/Stable classifier), permissively licensed, and has low install friction. However, maintenance is aging (535 days since last release), so it is best suited for established use cases where the algorithm is known to work well for your content type. Verify that lxml compiles on your target platform and that boilerplate detection accuracy meets your needs before production use.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • lxml requires libxml2 and libxslt system libraries; may need compilation on some platforms.
  • Low friction: pure Python wheel, only two runtime dependencies (lxml and backports.functools-lru-cache).
  • Maintenance status is aging—last release 535 days ago, but repository is active (last commit 2025-02-25) with no archived flag.

License · maintenance · safety

The BSD 2-Clause License (permissive) — BSD 2-Clause License (permissive). No restrictions on commercial use, modification, or redistribution; standard permissive terms apply.

last release 2025-02-25 (535 days) · last repo commit 2025-02-25 · 822 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 11,719,151 downloads/mo, #1,368 on PyPI

Verify before relying

import justext

paragraphs = justext.justext(html_content, justext.get_stoplist("English"))
for paragraph in paragraphs:
    if not paragraph.is_boilerplate:
        print(paragraph.text)
  • Performance characteristics on very large HTML documents or high-throughput scenarios
  • Accuracy of boilerplate detection across non-English languages and modern web layouts
  • Whether backports.functools-lru-cache is still required on modern Python versions
Same gist for agents: .md · .json

What it is and what it does

jusText is a heuristic-based tool for extracting main content from HTML pages by filtering out boilerplate elements like navigation, headers, and footers. It operates by analyzing text density and sentence structure, making it particularly suited for building linguistic corpora and web-scraped datasets where clean, sentence-bearing text is needed.

The package provides both a command-line interface (via `python -m justext`) and a Python API. It depends on lxml for HTML parsing and includes language-specific stoplists to improve detection accuracy. The core algorithm preserves paragraphs containing full sentences while marking others as boilerplate, allowing you to filter results programmatically.

Use it for

  • Build web corpora for natural language processing by extracting clean text from crawled HTML pages
  • Preprocess web content before feeding it into linguistic analysis or text processing pipelines
  • Automate extraction of article text from news sites or blogs while discarding navigation and ads
  • Clean HTML snapshots for archival or readability analysis in web preservation workflows
  • Remove noise from web-scraped training data for text classification or language models

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

The package is stable (Production/Stable classifier), permissively licensed, and has low install friction. However, maintenance is aging (535 days since last release), so it is best suited for established use cases where the algorithm is known to work well for your content type. Verify that lxml compiles on your target platform and that boilerplate detection accuracy meets your needs before production use.

Install

justext on PyPI

Before you install

Low friction: pure Python wheel, only two runtime dependencies (lxml and backports.functools-lru-cache). Maintenance status is aging—last release 535 days ago, but repository is active (last commit 2025-02-25) with no archived flag.

lxml requires libxml2 and libxslt system libraries; may need compilation on some platforms.

License in practice

BSD 2-Clause License (permissive). No restrictions on commercial use, modification, or redistribution; standard permissive terms apply.

Quickstart

import justext

paragraphs = justext.justext(html_content, justext.get_stoplist("English"))
for paragraph in paragraphs:
    if not paragraph.is_boilerplate:
        print(paragraph.text)

Verify before relying

  • Performance characteristics on very large HTML documents or high-throughput scenarios
  • Accuracy of boilerplate detection across non-English languages and modern web layouts
  • Whether backports.functools-lru-cache is still required on modern Python versions

Package facts

LicenseThe BSD 2-Clause License permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
lxmlbackports.functools-lru-cache
MaintenanceAging 535 days since the last release
Last repo commit
First released
Downloads11,719,151 / month, #1,368 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonTopic :: Internet :: WWW/HTTPTopic :: Software Development :: Pre-processorsTopic :: Text Processing :: FiltersTopic :: Text Processing :: Markup :: HTML

Evidence: justext-3.0.2-py2.py3-none-any.whl

Tags

Capabilities
html boilerplate removalextract main text from htmlweb content extractionremove navigation headers footershtml text cleaninglinguistic corpus buildingboilerplate stripping
Topics
content-extractionweb-scrapingnlp-preprocessing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “html boilerplate removal”

  • jusTextjusText removes boilerplate content (navigation, headers, footers)…
  • boilerpy3Extracts main article text and content from HTML, filtering out…
  • breadabilityExtracts the main readable content from HTML pages by identifying and…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also breadability · html-text · sumy · trafilatura · readability-lxml · inscriptis · MainContentExtractor · html2text · docx2txt · parsel