MainContentExtractor
A library to extract the main content from html. Developed for information on LLM and for feeding data into LangChain and LlamaIndex.
Decision gist · record as of 2026-08-14
Yes, if you need HTML-format output from main content extraction for LLM workflows. The library fills a real gap and has low install friction. However, proceed with caution: the package has been dormant since 2023-12-10 with no updates, maintenance status is unclear, and the XML-to-HTML conversion is lossy. Evaluate whether trafilatura alone or an alternative meets your needs before committing to a package with uncertain long-term support.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with three straightforward dependencies.
- Package is dormant since its single release on 2023-12-10, with no updates despite a recent commit on 2024-05-16; maintenance status is unclear and future updates are uncertain.
License · maintenance · safety
MIT (permissive) — MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
last release 2023-12-10 (978 days) · last repo commit 2024-05-16 · 51 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 79,301 downloads/mo, #14,367 on PyPI
Alternatives
Verify before relying
pip install MainContentExtractor
from main_content_extractor import MainContentExtractor
extracted_html = MainContentExtractor.extract(html_content)
extracted_markdown = MainContentExtractor.extract(html_content, output_format="markdown")- Whether the May 2024 commit represents ongoing development or a one-time fix after the December 2023 release.
- How the XML-to-HTML conversion affects data fidelity in practice and whether the irreversibility is a significant limitation.
- Whether output quality and format preservation meet requirements compared to using trafilatura directly.
What it is and what it does
MainContentExtractor wraps trafilatura to extract the main article or content body from HTML pages while preserving structural information like links and headers. Unlike trafilatura alone, it can output in HTML format (via an XML intermediate step), plain text, or Markdown—formats more suitable for feeding into LLM frameworks. It depends on trafilatura for core extraction, html2text for text conversion, and beautifulsoup4 for HTML parsing.
The library addresses a specific gap: trafilatura is effective at identifying main content but cannot output HTML directly and sometimes loses necessary data. MainContentExtractor compensates by converting trafilatura's XML output to HTML, though this conversion is not perfectly reversible. It's designed for workflows where you need both structured content and multiple output formats, particularly when preparing web data for language models.
Use it for
- Prepare web articles for ingestion into LangChain or LlamaIndex by extracting clean content in Markdown or text format.
- Extract main content from HTML while preserving link and header hierarchy for downstream analysis or indexing.
- Convert web pages to Markdown for easier consumption by LLM prompts or knowledge bases.
- Remove boilerplate, navigation, and ads from web pages before processing with NLP tools.
- Build web scraping pipelines that output multiple formats from a single extraction pass.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need HTML-format output from main content extraction for LLM workflows.
The library fills a real gap and has low install friction. However, proceed with caution: the package has been dormant since 2023-12-10 with no updates, maintenance status is unclear, and the XML-to-HTML conversion is lossy. Evaluate whether trafilatura alone or an alternative meets your needs before committing to a package with uncertain long-term support.
Install
maincontentextractor on PyPI
Before you install
Low install friction with three straightforward dependencies. Package is dormant since its single release on 2023-12-10, with no updates despite a recent commit on 2024-05-16; maintenance status is unclear and future updates are uncertain.
License in practice
MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
Quickstart
pip install MainContentExtractor
from main_content_extractor import MainContentExtractor
extracted_html = MainContentExtractor.extract(html_content)
extracted_markdown = MainContentExtractor.extract(html_content, output_format="markdown")
Verify before relying
- Whether the May 2024 commit represents ongoing development or a one-time fix after the December 2023 release.
- How the XML-to-HTML conversion affects data fidelity in practice and whether the irreversibility is a significant limitation.
- Whether output quality and format preservation meet requirements compared to using trafilatura directly.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.6 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagestrafilaturahtml2textbeautifulsoup4 |
| Maintenance | Dormant 978 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 79,301 / month, #14,367 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: MainContentExtractor-0.0.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “extract main content from html”
- MainContentExtractorExtracts main content from HTML pages and outputs it in HTML, text,…
- readability-lxmlExtracts and cleans the main article text and title from HTML…
- readable-contentExtracts the main article content from web pages, handling both…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also trafilatura · jusText · pymupdf4llm · langextract · semantic-text-splitter · Crawl4AI · langchain-tavily · inscriptis · html-text · pdfminer