warc3-wet-clueweb09
Python library to work with ARC and WARC files, with fixes for ClueWeb09
Decision gist · record as of 2026-08-14
No, unless you have a specific, narrow need to read ClueWeb09 archives and cannot use an actively maintained alternative. The package is abandoned, Python version support is unspecified, and GPLv2 copyleft licensing restricts use in proprietary projects. High install friction and no maintenance path make it a poor choice for new projects.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Python version support is unspecified; verify compatibility with your Python version before relying on this abandoned package.
- Installation requires building from source with no runtime dependencies, which adds friction.
- The package is abandoned (last release 2020-12-07, no commits since 2021-12-06), so expect no maintenance or bug fixes going forward.
License · maintenance · safety
GPLv2 (copyleft) — Licensed under GPLv2 (copyleft), which requires that any derivative work or distribution must also be open-source under compatible terms—a significant constraint for proprietary or closed-source projects.
last release 2020-12-07 (2076 days) · last repo commit 2021-12-06
0 known vulnerabilities (OSV.dev, 2026-08-14) · 142,309 downloads/mo, #11,213 on PyPI
Alternatives
Verify before relying
pip install warc3-wet-clueweb09
with warc3_wet_clueweb09.open("test.warc") as f:
for record in f:
print(record['WARC-Target-URI'], record['Content-Length'])- Whether the package actually works with modern Python versions (support unspecified in metadata).
- Extent of ClueWeb09-specific fixes and whether they address all known parsing issues.
- Whether seeking support (mentioned in changelog for 0.2.3) is fully functional.
What it is and what it does
warc3-wet-clueweb09 is a Python library for reading WARC (Web ARChive) and WET (WARC Extracted Text) files, which are standard formats for storing web crawls. It provides an interface to iterate over archive records and access their metadata fields and content. The package is a Python 3 port of an older library, with modifications to handle parsing issues specific to ClueWeb09 archives.
The library has no runtime dependencies and is designed for straightforward use: open an archive file, iterate over records, and extract fields like WARC-Target-URI and Content-Length. However, the package is abandoned (last updated 2020-12-07 with no commits since 2021-12-06), so it receives no maintenance, bug fixes, or updates for compatibility with newer Python versions.
Use it for
- Parse ClueWeb09 web crawl archives to extract URLs and content metadata for research or analysis.
- Iterate over WARC files to build indexes or extract specific records for information retrieval tasks.
- Read WET files to access extracted text content from archived web pages.
- Migrate legacy Python 2 WARC processing pipelines to Python 3 using this ported library.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No, unless you have a specific, narrow need to read ClueWeb09 archives and cannot use an actively maintained alternative.
The package is abandoned, Python version support is unspecified, and GPLv2 copyleft licensing restricts use in proprietary projects. High install friction and no maintenance path make it a poor choice for new projects.
Install
warc3-wet-clueweb09 on PyPI
Before you install
Installation requires building from source with no runtime dependencies, which adds friction. The package is abandoned (last release 2020-12-07, no commits since 2021-12-06), so expect no maintenance or bug fixes going forward.
Python version support is unspecified; verify compatibility with your Python version before relying on this abandoned package.
License in practice
Licensed under GPLv2 (copyleft), which requires that any derivative work or distribution must also be open-source under compatible terms—a significant constraint for proprietary or closed-source projects.
Quickstart
pip install warc3-wet-clueweb09
with warc3_wet_clueweb09.open("test.warc") as f:
for record in f:
print(record['WARC-Target-URI'], record['Content-Length'])
Verify before relying
- Whether the package actually works with modern Python versions (support unspecified in metadata).
- Extent of ClueWeb09-specific fixes and whether they address all known parsing issues.
- Whether seeking support (mentioned in changelog for 0.2.3) is fully functional.
Package facts
| License | GPLv2 copyleft |
| Python support | Not specified |
| Install friction | High. Source build required |
| Runtime dependencies | None |
| Maintenance | Abandoned 2,076 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 142,309 / month, #11,213 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: Web EnvironmentIntended Audience :: DevelopersLicense :: OSI Approved :: GNU General Public License v2 (GPLv2)Operating System :: OS IndependentProgramming Language :: Python |
Evidence: warc3-wet-clueweb09-0.2.5.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “warc file reader python”
- warc3-wet-clueweb09Reads and iterates over WARC and WET archive files, extracting…
- warc3-wetRead and iterate over WARC (Web ARChive) and WET (WARC Extracted…
- FastWARCFastWARC parses WARC (Web ARChive) files—compressed or…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also warc3-wet · warcio · FastWARC · Resiliparse · Crawl4AI · rarfile · extractcode · patool · pymarc · internetarchive