$npx skillfedfor your agent

warc3-wet

Python library to work with ARC and WARC files

With conditionsPyPI WWW/HTTPReleased Jul 2024219.9K downloads / moGPLv2Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — warc3_wet-0.2.5-py3-none-any.whl
v0.2.5 · released 2024-07-17

Yes, if you need to work with WARC or WET files. The package is stable, has no dependencies, and handles a specific file format well. The dormant maintenance status is not a blocker for read-only use, but verify GPLv2 compatibility with your project first. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with no runtime dependencies.
  • Maintenance is dormant—last release was 758 days ago—but the repository remains active.

License · maintenance · safety

GPLv2 (copyleft) — GPLv2 copyleft license means any derivative work or distribution must also be licensed under GPLv2 and have source code available; verify compatibility with your project's licensing before use.

last release 2024-07-17 (758 days) · last repo commit 2024-07-17 · 9 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 219,855 downloads/mo, #9,319 on PyPI

Verify before relying

pip install warc3-wet

with warc3_wet.open("test.warc") as f:
    for record in f:
        print(record['WARC-Target-URI'], record['Content-Length'])
  • Whether Python 3.10+ is fully supported (changelog mentions 3.10 compatibility in 0.2.4, but requires_python is unspecified)
  • Current stability and whether dormant status indicates maintenance-only mode or active development
Same gist for agents: .md · .json

What it is and what it does

warc3-wet is a Python library for reading and parsing WARC and WET archive files—the standard formats used to store web crawls. It provides a simple interface to open these files and iterate over records, extracting metadata headers like WARC-Target-URI and Content-Length. The library is a Python 3 port of an older warc package; this fork adds support for WET files and seeking within archives.

The package has no external runtime dependencies, making it lightweight to install. It is classified as Beta-status software. The repository remains accessible but has not seen a release in over two years, suggesting the package is in a stable, maintenance-only state.

Use it for

  • Extract and process records from web crawl archives stored in WARC format for data analysis or research.
  • Parse WET files to access extracted text content from archived web pages.
  • Build data pipelines that read WARC archives and filter or transform records based on metadata.
  • Access specific records within large WARC files using seek functionality for efficient random access.
  • Integrate web archive data into workflows that require crawled web content.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to work with WARC or WET files.

The package is stable, has no dependencies, and handles a specific file format well. The dormant maintenance status is not a blocker for read-only use, but verify GPLv2 compatibility with your project first. No known security vulnerabilities.

Install

warc3-wet on PyPI

Before you install

Low install friction with no runtime dependencies. Maintenance is dormant—last release was 758 days ago—but the repository remains active.

License in practice

GPLv2 copyleft license means any derivative work or distribution must also be licensed under GPLv2 and have source code available; verify compatibility with your project's licensing before use.

Quickstart

pip install warc3-wet

with warc3_wet.open("test.warc") as f:
    for record in f:
        print(record['WARC-Target-URI'], record['Content-Length'])

Verify before relying

  • Whether Python 3.10+ is fully supported (changelog mentions 3.10 compatibility in 0.2.4, but requires_python is unspecified)
  • Current stability and whether dormant status indicates maintenance-only mode or active development

Package facts

LicenseGPLv2 copyleft
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceDormant 758 days since the last release
Last repo commit
First released
Downloads219,855 / month, #9,319 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: Web EnvironmentIntended Audience :: DevelopersLicense :: OSI Approved :: GNU General Public License v2 (GPLv2)Operating System :: OS IndependentProgramming Language :: Python

Evidence: warc3_wet-0.2.5-py3-none-any.whl

Tags

Capabilities
warc file parserweb archive readerwarc format librarywet file parsingcrawl archive extractionwarc record iterationweb crawl data access
Topics
web-archivingdata-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “warc file parser”

  • warc3-wetRead and iterate over WARC (Web ARChive) and WET (WARC Extracted…
  • warc3-wet-clueweb09Reads and iterates over WARC and WET archive files, extracting…
  • FastWARCFastWARC parses WARC (Web ARChive) files—compressed or…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also warc3-wet-clueweb09 · warcio · FastWARC · Resiliparse · rarfile · Crawl4AI · remotezip · patool · extractcode · internetarchive