trafilatura
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.
Decision gist · record as of 2026-08-14
Yes. Trafilatura is actively maintained, widely adopted, has no known vulnerabilities, and offers low install friction. It solves a real problem (converting messy HTML to clean text and metadata) with a mature, well-documented API. Use it if you need reliable web content extraction for research, data pipelines, or corpus building.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- lxml (a runtime dependency) requires compilation; pre-built wheels are available for most platforms.
- Low install friction with a pure-Python wheel distribution.
License · maintenance · safety
Apache-2.0 (permissive) — Distributed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions. Versions prior to 1.8.0 were GPLv3+, but current releases impose no copyleft obligations.
last release 2026-07-31 (14 days) · last repo commit 2026-08-14 · 6,633 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 13,364,051 downloads/mo, #1,284 on PyPI
Alternatives
Verify before relying
pip install trafilatura
import trafilatura
downloaded = trafilatura.fetch_url('https://example.com')
result = trafilatura.extract(downloaded)
print(result)- Whether the package's performance benchmarks (ScrapingHub, Bevendorff et al. 2023) remain current for version 2.2.0.
- Specific memory or CPU constraints when processing very large HTML documents or high-volume crawls.
- Whether language detection and speed optimization add-ons are included or require separate installation.
What it is and what it does
Trafilatura is a Python library and command-line tool for discovering, downloading, and extracting text and metadata from web pages. It handles the full pipeline: crawling via sitemaps and feeds, downloading HTML, parsing and cleaning content, and exporting results in formats like JSON, CSV, Markdown, XML, and HTML. The core extraction logic balances precision (removing noise like headers, footers, navigation) against recall (preserving valid content), and it supports optional extraction of comments, links, images, and tables alongside main text and metadata.
The package is designed for both one-off extraction and bulk processing, with parallel handling of online URLs and offline HTML files. It integrates widely in academic and commercial projects (HuggingFace, IBM, Microsoft Research, Stanford, Allen Institute) and is actively maintained with regular updates. Runtime dependencies include lxml for parsing, courlan for URL management, htmldate for date extraction, and justext for text segmentation.
Use it for
- Build a text corpus from news websites or blogs by crawling sitemaps and extracting article text and publication dates.
- Convert downloaded HTML files to clean JSON or CSV for data analysis, removing boilerplate and preserving only article content.
- Extract metadata (title, author, date, categories) from web pages for indexing or cataloging.
- Process large batches of HTML documents in parallel to prepare training data for NLP or machine learning models.
- Scrape and structure web content for research purposes while respecting politeness and deduplication rules.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Trafilatura is actively maintained, widely adopted, has no known vulnerabilities, and offers low install friction. It solves a real problem (converting messy HTML to clean text and metadata) with a mature, well-documented API. Use it if you need reliable web content extraction for research, data pipelines, or corpus building.
Install
trafilatura on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release (14 days old) and strong community engagement (6633 GitHub stars). Supports Python 3.10 through 3.14.
Requires Python 3.10 or later. lxml (a runtime dependency) requires compilation; pre-built wheels are available for most platforms.
License in practice
Distributed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions. Versions prior to 1.8.0 were GPLv3+, but current releases impose no copyleft obligations.
Quickstart
pip install trafilatura
import trafilatura
downloaded = trafilatura.fetch_url('https://example.com')
result = trafilatura.extract(downloaded)
print(result)
Verify before relying
- Whether the package's performance benchmarks (ScrapingHub, Bevendorff et al. 2023) remain current for version 2.2.0.
- Specific memory or CPU constraints when processing very large HTML documents or high-volume crawls.
- Whether language detection and speed optimization add-ons are included or require separate installation.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagescertificharset_normalizercourlanhtmldatejustextlxmlurllib3 |
| Maintenance | Actively maintained 14 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 13,364,051 / month, #1,284 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOSOperating System :: MicrosoftOperating System :: POSIXProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTPTopic :: Scientific/Engineering :: Information AnalysisTopic :: SecurityTopic :: Text Editors :: Text ProcessingTopic :: Text Processing :: LinguisticTopic :: Text Processing :: Markup :: HTMLTopic :: Text Processing :: Markup :: MarkdownTopic :: Text Processing :: Markup :: XMLTopic :: Utilities |
Evidence: trafilatura-2.2.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “web crawling and text mining”
- trafilaturaTrafilatura extracts main text, metadata, and structured content from…
- LinkCheckerLinkChecker validates links across websites by crawling pages and…
- newspaper3kNewspaper3k downloads and extracts article text, metadata, images,…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also htmldate · MainContentExtractor · jusText · readable-content · Scrapy · courlan · Crawl4AI · scholarly · boilerpy3 · newspaper4k