FastWARC
The world's fastest WARC parsing library written in Rust with bindings for Python.
Decision gist · record as of 2026-08-14
Yes. FastWARC is actively maintained, has no known vulnerabilities, is licensed permissively, and offers pre-built wheels for modern Python versions on all major platforms. Install it if you work with WARC files or web archives; the Rust backend and broad compression support make it a solid choice for this specific task.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; pre-built wheels available for macOS (arm64, x86_64), Linux (aarch64, x86_64), and Windows (amd64).
- Pre-built wheels cover modern Python versions and major platforms (macOS, Linux, Windows), reducing compile friction.
- The package is actively maintained with a recent release and no known vulnerabilities.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 is permissive; you may use, modify, and distribute FastWARC freely in proprietary or open-source projects with minimal restriction.
last release 2026-07-20 (25 days) · last repo commit 2026-07-20 · 144 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,402,572 downloads/mo, #3,945 on PyPI
Alternatives
Verify before relying
pip install fastwarc
from fastwarc.warc import ArchiveIterator
with open('archive.warc.gz', 'rb') as f:
for record in ArchiveIterator(f):
print(record.headers)- Exact performance improvement over pure-Python WARC parsers is not quantified in the fact sheet.
- Memory overhead or streaming behavior for very large WARC files is not documented in the excerpt.
- Whether the library supports incremental parsing or requires loading entire records into memory.
What it is and what it does
FastWARC is a Python library that reads WARC archive files—the standard format for storing web crawl data—using a Rust implementation for speed. It handles both compressed (Gzip, Zstd, LZ4) and uncompressed WARC/1.0 and WARC/1.1 streams. The library is part of the ChatNoir Resiliparse toolkit for web data processing and is designed for workloads that need to ingest and parse large volumes of web archive data efficiently.
You install it via pip and import it to iterate over records in WARC files. It depends on brotli, click, tqdm, and typing_extensions. The package requires Python 3.10 or later and is actively maintained with recent releases and no known security issues.
Use it for
- Parse Common Crawl WARC dumps or other web archive collections for research or data extraction.
- Build web crawl pipelines that need to read compressed WARC files without decompressing to disk first.
- Extract HTTP responses, metadata, or payloads from archived web data at scale.
- Integrate WARC parsing into information retrieval or web analytics workflows.
- Process historical web snapshots stored in WARC format for archival or compliance tasks.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
FastWARC is actively maintained, has no known vulnerabilities, is licensed permissively, and offers pre-built wheels for modern Python versions on all major platforms. Install it if you work with WARC files or web archives; the Rust backend and broad compression support make it a solid choice for this specific task.
Install
fastwarc on PyPI
Before you install
Pre-built wheels cover modern Python versions and major platforms (macOS, Linux, Windows), reducing compile friction. The package is actively maintained with a recent release and no known vulnerabilities.
Requires Python 3.10 or later; pre-built wheels available for macOS (arm64, x86_64), Linux (aarch64, x86_64), and Windows (amd64).
License in practice
Apache-2.0 is permissive; you may use, modify, and distribute FastWARC freely in proprietary or open-source projects with minimal restriction.
Quickstart
pip install fastwarc
from fastwarc.warc import ArchiveIterator
with open('archive.warc.gz', 'rb') as f:
for record in ArchiveIterator(f):
print(record.headers)
Verify before relying
- Exact performance improvement over pure-Python WARC parsers is not quantified in the fact sheet.
- Memory overhead or streaming behavior for very large WARC files is not documented in the excerpt.
- Whether the library supports incremental parsing or requires loading entire records into memory.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 4 packagesbrotliclicktqdmtyping_extensions |
| Maintenance | Actively maintained 25 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,402,572 / month, #3,945 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: fastwarc-1.0.9-cp310-cp310-macosx_12_0_arm64.whl; fastwarc-1.0.9-cp310-cp310-macosx_12_0_x86_64.whl; fastwarc-1.0.9-cp310-cp310-manylinux_2_28_aarch64.whl; fastwarc-1.0.9-cp310-cp310-manylinux_2_28_x86_64.whl; fastwarc-1.0.9-cp310-cp310-win_amd64.whl; fastwarc-1.0.9-cp311-cp311-macosx_12_0_arm64.whl; fastwarc-1.0.9-cp311-cp311-macosx_12_0_x86_64.whl; fastwarc-1.0.9-cp311-cp311-manylinux_2_28_aarch64.whl; fastwarc-1.0.9-cp311-cp311-manylinux_2_28_x86_64.whl; fastwarc-1.0.9-cp311-cp311-win_amd64.whl; fastwarc-1.0.9-cp312-cp312-macosx_12_0_arm64.whl; fastwarc-1.0.9-cp312-cp312-macosx_12_0_x86_64.whl; fastwarc-1.0.9-cp312-cp312-manylinux_2_28_aarch64.whl; fastwarc-1.0.9-cp312-cp312-manylinux_2_28_x86_64.whl; fastwarc-1.0.9-cp312-cp312-win_amd64.whl; fastwarc-1.0.9-cp313-cp313-macosx_12_0_arm64.whl; fastwarc-1.0.9-cp313-cp313-macosx_12_0_x86_64.whl; fastwarc-1.0.9-cp313-cp313-manylinux_2_28_aarch64.whl; fastwarc-1.0.9-cp313-cp313-manylinux_2_28_x86_64.whl; fastwarc-1.0.9-cp313-cp313-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “warc file parsing”
- FastWARCFastWARC parses WARC (Web ARChive) files—compressed or…
- warc3-wet-clueweb09Reads and iterates over WARC and WET archive files, extracting…
- warc3-wetRead and iterate over WARC (Web ARChive) and WET (WARC Extracted…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also Resiliparse · warcio · warc3-wet · warc3-wet-clueweb09 · fastar · cramjam · gzip-stream · Flask-Compress · python-lzf · ar