warc3-wet
Python library to work with ARC and WARC files
Decision gist · record as of 2026-08-14
Yes, if you need to work with WARC or WET files. The package is stable, has no dependencies, and handles a specific file format well. The dormant maintenance status is not a blocker for read-only use, but verify GPLv2 compatibility with your project first. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with no runtime dependencies.
- Maintenance is dormant—last release was 758 days ago—but the repository remains active.
License · maintenance · safety
GPLv2 (copyleft) — GPLv2 copyleft license means any derivative work or distribution must also be licensed under GPLv2 and have source code available; verify compatibility with your project's licensing before use.
last release 2024-07-17 (758 days) · last repo commit 2024-07-17 · 9 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 219,855 downloads/mo, #9,319 on PyPI
Alternatives
Verify before relying
pip install warc3-wet
with warc3_wet.open("test.warc") as f:
for record in f:
print(record['WARC-Target-URI'], record['Content-Length'])- Whether Python 3.10+ is fully supported (changelog mentions 3.10 compatibility in 0.2.4, but requires_python is unspecified)
- Current stability and whether dormant status indicates maintenance-only mode or active development
What it is and what it does
warc3-wet is a Python library for reading and parsing WARC and WET archive files—the standard formats used to store web crawls. It provides a simple interface to open these files and iterate over records, extracting metadata headers like WARC-Target-URI and Content-Length. The library is a Python 3 port of an older warc package; this fork adds support for WET files and seeking within archives.
The package has no external runtime dependencies, making it lightweight to install. It is classified as Beta-status software. The repository remains accessible but has not seen a release in over two years, suggesting the package is in a stable, maintenance-only state.
Use it for
- Extract and process records from web crawl archives stored in WARC format for data analysis or research.
- Parse WET files to access extracted text content from archived web pages.
- Build data pipelines that read WARC archives and filter or transform records based on metadata.
- Access specific records within large WARC files using seek functionality for efficient random access.
- Integrate web archive data into workflows that require crawled web content.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to work with WARC or WET files.
The package is stable, has no dependencies, and handles a specific file format well. The dormant maintenance status is not a blocker for read-only use, but verify GPLv2 compatibility with your project first. No known security vulnerabilities.
Install
warc3-wet on PyPI
Before you install
Low install friction with no runtime dependencies. Maintenance is dormant—last release was 758 days ago—but the repository remains active.
License in practice
GPLv2 copyleft license means any derivative work or distribution must also be licensed under GPLv2 and have source code available; verify compatibility with your project's licensing before use.
Quickstart
pip install warc3-wet
with warc3_wet.open("test.warc") as f:
for record in f:
print(record['WARC-Target-URI'], record['Content-Length'])
Verify before relying
- Whether Python 3.10+ is fully supported (changelog mentions 3.10 compatibility in 0.2.4, but requires_python is unspecified)
- Current stability and whether dormant status indicates maintenance-only mode or active development
Package facts
| License | GPLv2 copyleft |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Dormant 758 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 219,855 / month, #9,319 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: Web EnvironmentIntended Audience :: DevelopersLicense :: OSI Approved :: GNU General Public License v2 (GPLv2)Operating System :: OS IndependentProgramming Language :: Python |
Evidence: warc3_wet-0.2.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “warc file parser”
- warc3-wetRead and iterate over WARC (Web ARChive) and WET (WARC Extracted…
- warc3-wet-clueweb09Reads and iterates over WARC and WET archive files, extracting…
- FastWARCFastWARC parses WARC (Web ARChive) files—compressed or…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also warc3-wet-clueweb09 · warcio · FastWARC · Resiliparse · rarfile · Crawl4AI · remotezip · patool · extractcode · internetarchive