$npx skillfedfor your agent

trafilatura

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.

Worth itPyPI WWW/HTTPReleased Jul 202613.4M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — trafilatura-2.2.0-py3-none-any.whl
v2.2.0 · released 2026-07-31 · Python >=3.10 · 7 runtime deps: certifi, charset_normalizer, courlan, htmldate, justext, lxml, urllib3

Yes. Trafilatura is actively maintained, widely adopted, has no known vulnerabilities, and offers low install friction. It solves a real problem (converting messy HTML to clean text and metadata) with a mature, well-documented API. Use it if you need reliable web content extraction for research, data pipelines, or corpus building.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • lxml (a runtime dependency) requires compilation; pre-built wheels are available for most platforms.
  • Low install friction with a pure-Python wheel distribution.

License · maintenance · safety

Apache-2.0 (permissive) — Distributed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions. Versions prior to 1.8.0 were GPLv3+, but current releases impose no copyleft obligations.

last release 2026-07-31 (14 days) · last repo commit 2026-08-14 · 6,633 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 13,364,051 downloads/mo, #1,284 on PyPI

Verify before relying

pip install trafilatura

import trafilatura
downloaded = trafilatura.fetch_url('https://example.com')
result = trafilatura.extract(downloaded)
print(result)
  • Whether the package's performance benchmarks (ScrapingHub, Bevendorff et al. 2023) remain current for version 2.2.0.
  • Specific memory or CPU constraints when processing very large HTML documents or high-volume crawls.
  • Whether language detection and speed optimization add-ons are included or require separate installation.
Same gist for agents: .md · .json

What it is and what it does

Trafilatura is a Python library and command-line tool for discovering, downloading, and extracting text and metadata from web pages. It handles the full pipeline: crawling via sitemaps and feeds, downloading HTML, parsing and cleaning content, and exporting results in formats like JSON, CSV, Markdown, XML, and HTML. The core extraction logic balances precision (removing noise like headers, footers, navigation) against recall (preserving valid content), and it supports optional extraction of comments, links, images, and tables alongside main text and metadata.

The package is designed for both one-off extraction and bulk processing, with parallel handling of online URLs and offline HTML files. It integrates widely in academic and commercial projects (HuggingFace, IBM, Microsoft Research, Stanford, Allen Institute) and is actively maintained with regular updates. Runtime dependencies include lxml for parsing, courlan for URL management, htmldate for date extraction, and justext for text segmentation.

Use it for

  • Build a text corpus from news websites or blogs by crawling sitemaps and extracting article text and publication dates.
  • Convert downloaded HTML files to clean JSON or CSV for data analysis, removing boilerplate and preserving only article content.
  • Extract metadata (title, author, date, categories) from web pages for indexing or cataloging.
  • Process large batches of HTML documents in parallel to prepare training data for NLP or machine learning models.
  • Scrape and structure web content for research purposes while respecting politeness and deduplication rules.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Trafilatura is actively maintained, widely adopted, has no known vulnerabilities, and offers low install friction. It solves a real problem (converting messy HTML to clean text and metadata) with a mature, well-documented API. Use it if you need reliable web content extraction for research, data pipelines, or corpus building.

Install

trafilatura on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release (14 days old) and strong community engagement (6633 GitHub stars). Supports Python 3.10 through 3.14.

Requires Python 3.10 or later. lxml (a runtime dependency) requires compilation; pre-built wheels are available for most platforms.

License in practice

Distributed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions. Versions prior to 1.8.0 were GPLv3+, but current releases impose no copyleft obligations.

Quickstart

pip install trafilatura

import trafilatura
downloaded = trafilatura.fetch_url('https://example.com')
result = trafilatura.extract(downloaded)
print(result)

Verify before relying

  • Whether the package's performance benchmarks (ScrapingHub, Bevendorff et al. 2023) remain current for version 2.2.0.
  • Specific memory or CPU constraints when processing very large HTML documents or high-volume crawls.
  • Whether language detection and speed optimization add-ons are included or require separate installation.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
certificharset_normalizercourlanhtmldatejustextlxmlurllib3
MaintenanceActively maintained 14 days since the last release
Last repo commit
First released
Downloads13,364,051 / month, #1,284 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOSOperating System :: MicrosoftOperating System :: POSIXProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTPTopic :: Scientific/Engineering :: Information AnalysisTopic :: SecurityTopic :: Text Editors :: Text ProcessingTopic :: Text Processing :: LinguisticTopic :: Text Processing :: Markup :: HTMLTopic :: Text Processing :: Markup :: MarkdownTopic :: Text Processing :: Markup :: XMLTopic :: Utilities

Evidence: trafilatura-2.2.0-py3-none-any.whl

Tags

Capabilities
web scraping text extractionhtml to text converterweb content extractionmetadata extraction from htmlnews article scraperweb crawling and text mininghtml parsing and cleaning
Topics
web-scrapingtext-extractionnlp-data-prep
PyPI keywords
corpushtml2textnews-crawlernatural-language-processingscrapertei-xmltext-extractionwebscrapingweb-scraping

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “web crawling and text mining”

  • trafilaturaTrafilatura extracts main text, metadata, and structured content from…
  • LinkCheckerLinkChecker validates links across websites by crawling pages and…
  • newspaper3kNewspaper3k downloads and extracts article text, metadata, images,…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also htmldate · MainContentExtractor · jusText · readable-content · Scrapy · courlan · Crawl4AI · scholarly · boilerpy3 · newspaper4k