$npx skillfedfor your agent

MainContentExtractor

A library to extract the main content from html. Developed for information on LLM and for feeding data into LangChain and LlamaIndex.

With conditionsPyPI HTMLReleased Dec 202379.3K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — MainContentExtractor-0.0.4-py3-none-any.whl
v0.0.4 · released 2023-12-10 · Python >=3.6 · 3 runtime deps: trafilatura, html2text, beautifulsoup4

Yes, if you need HTML-format output from main content extraction for LLM workflows. The library fills a real gap and has low install friction. However, proceed with caution: the package has been dormant since 2023-12-10 with no updates, maintenance status is unclear, and the XML-to-HTML conversion is lossy. Evaluate whether trafilatura alone or an alternative meets your needs before committing to a package with uncertain long-term support.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with three straightforward dependencies.
  • Package is dormant since its single release on 2023-12-10, with no updates despite a recent commit on 2024-05-16; maintenance status is unclear and future updates are uncertain.

License · maintenance · safety

MIT (permissive) — MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

last release 2023-12-10 (978 days) · last repo commit 2024-05-16 · 51 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 79,301 downloads/mo, #14,367 on PyPI

Verify before relying

pip install MainContentExtractor

from main_content_extractor import MainContentExtractor
extracted_html = MainContentExtractor.extract(html_content)
extracted_markdown = MainContentExtractor.extract(html_content, output_format="markdown")
  • Whether the May 2024 commit represents ongoing development or a one-time fix after the December 2023 release.
  • How the XML-to-HTML conversion affects data fidelity in practice and whether the irreversibility is a significant limitation.
  • Whether output quality and format preservation meet requirements compared to using trafilatura directly.
Same gist for agents: .md · .json

What it is and what it does

MainContentExtractor wraps trafilatura to extract the main article or content body from HTML pages while preserving structural information like links and headers. Unlike trafilatura alone, it can output in HTML format (via an XML intermediate step), plain text, or Markdown—formats more suitable for feeding into LLM frameworks. It depends on trafilatura for core extraction, html2text for text conversion, and beautifulsoup4 for HTML parsing.

The library addresses a specific gap: trafilatura is effective at identifying main content but cannot output HTML directly and sometimes loses necessary data. MainContentExtractor compensates by converting trafilatura's XML output to HTML, though this conversion is not perfectly reversible. It's designed for workflows where you need both structured content and multiple output formats, particularly when preparing web data for language models.

Use it for

  • Prepare web articles for ingestion into LangChain or LlamaIndex by extracting clean content in Markdown or text format.
  • Extract main content from HTML while preserving link and header hierarchy for downstream analysis or indexing.
  • Convert web pages to Markdown for easier consumption by LLM prompts or knowledge bases.
  • Remove boilerplate, navigation, and ads from web pages before processing with NLP tools.
  • Build web scraping pipelines that output multiple formats from a single extraction pass.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need HTML-format output from main content extraction for LLM workflows.

The library fills a real gap and has low install friction. However, proceed with caution: the package has been dormant since 2023-12-10 with no updates, maintenance status is unclear, and the XML-to-HTML conversion is lossy. Evaluate whether trafilatura alone or an alternative meets your needs before committing to a package with uncertain long-term support.

Install

maincontentextractor on PyPI

Before you install

Low install friction with three straightforward dependencies. Package is dormant since its single release on 2023-12-10, with no updates despite a recent commit on 2024-05-16; maintenance status is unclear and future updates are uncertain.

License in practice

MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

Quickstart

pip install MainContentExtractor

from main_content_extractor import MainContentExtractor
extracted_html = MainContentExtractor.extract(html_content)
extracted_markdown = MainContentExtractor.extract(html_content, output_format="markdown")

Verify before relying

  • Whether the May 2024 commit represents ongoing development or a one-time fix after the December 2023 release.
  • How the XML-to-HTML conversion affects data fidelity in practice and whether the irreversibility is a significant limitation.
  • Whether output quality and format preservation meet requirements compared to using trafilatura directly.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.6
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
trafilaturahtml2textbeautifulsoup4
MaintenanceDormant 978 days since the last release
Last repo commit
First released
Downloads79,301 / month, #14,367 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: MainContentExtractor-0.0.4-py3-none-any.whl

Tags

Capabilities
extract main content from htmlhtml content extractionremove boilerplate from webpageprepare html for llmconvert html to markdownweb scraping content extractionhtml to text conversion
Topics
web-scrapingllm-data-prepcontent-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract main content from html”

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also trafilatura · jusText · pymupdf4llm · langextract · semantic-text-splitter · Crawl4AI · langchain-tavily · inscriptis · html-text · pdfminer