internetarchive
A Python interface to archive.org.
Decision gist · record as of 2026-08-14
Yes, with a critical caveat: upgrade to v5.4.2+ immediately if you use File.download(), as versions <=5.4.1 contain a directory traversal vulnerability. The package is actively maintained, has low install friction, and is the standard way to access Archive.org programmatically. AGPL-3.0 licensing requires review if you plan proprietary use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10.
- Versions <=5.4.1 contain a critical directory traversal vulnerability in File.download(); upgrade to v5.4.2+ immediately.
- Low install friction with four lightweight runtime dependencies.
License · maintenance · safety
AGPL-3.0 (agpl) — Licensed under AGPL-3.0, which requires derivative works and modifications to be distributed under the same license. Suitable for open-source projects but may restrict commercial or proprietary use without careful licensing review.
last release 2026-07-22 (23 days) · last repo commit 2026-07-24 · 1,897 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 320,706 downloads/mo, #7,634 on PyPI
Alternatives
Verify before relying
pip install internetarchive
from internetarchive import get_item
item = get_item('example_item')
item.download()- Specific Archive.org API endpoints and rate limits supported by this version
- Whether authentication (Archive.org credentials) is required for all operations or only some
- Performance characteristics for bulk downloads or large collections
What it is and what it does
internetarchive is a Python library and command-line tool that provides access to Archive.org's digital collections. It allows developers to programmatically search for items, retrieve metadata, and download files from the Internet Archive, as well as interact with the service via the `ia` command-line tool. The package wraps Archive.org's APIs and handles the HTTP communication through requests and urllib3, with jsonpatch for metadata manipulation and tqdm for progress reporting during downloads.
The library is production-stable and actively maintained, with support for modern Python versions (3.10+). It is commonly installed via pipx for command-line use or pip for library integration. The package has a long release history since 2013 and maintains a healthy repository with recent commits, making it a reliable choice for Archive.org integration.
Use it for
- Download archived web pages, books, or media collections from Archive.org in bulk for research or preservation
- Build scripts that query Archive.org metadata to find and retrieve specific items programmatically
- Integrate Archive.org search and retrieval into Python applications without writing raw HTTP requests
- Use the `ia` command-line tool for one-off downloads or administrative tasks from the shell
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with a critical caveat: upgrade to v5.4.2+ immediately if you use File.download(), as versions <=5.4.1 contain a directory traversal vulnerability.
The package is actively maintained, has low install friction, and is the standard way to access Archive.org programmatically. AGPL-3.0 licensing requires review if you plan proprietary use.
Install
internetarchive on PyPI
Before you install
Low install friction with four lightweight runtime dependencies. Active maintenance: last release 23 days ago, repository last commit 2026-07-24, and 1897 stars indicate ongoing support. Requires Python >=3.10.
Requires Python >=3.10. Versions <=5.4.1 contain a critical directory traversal vulnerability in File.download(); upgrade to v5.4.2+ immediately.
License in practice
Licensed under AGPL-3.0, which requires derivative works and modifications to be distributed under the same license. Suitable for open-source projects but may restrict commercial or proprietary use without careful licensing review.
Quickstart
pip install internetarchive
from internetarchive import get_item
item = get_item('example_item')
item.download()
Verify before relying
- Specific Archive.org API endpoints and rate limits supported by this version
- Whether authentication (Archive.org credentials) is required for all operations or only some
- Performance characteristics for bulk downloads or large collections
Package facts
| License | AGPL-3.0 agpl |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesjsonpatchrequeststqdmurllib3 |
| Maintenance | Actively maintained 23 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 320,706 / month, #7,634 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersNatural Language :: EnglishProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy |
Evidence: internetarchive-5.11.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “archive.org api client”
- internetarchiveProgrammatic and command-line interface to search, browse, and…
- savepagenowWrapper and CLI tool to submit URLs to archive.org's Save Page Now…
- codewords-clientA Python client library for the Codewords API with built-in FastAPI…
Give your agent the search over MCP, or paste the wish link into any chat.
More Internet packages
Botocore provides low-level, data-driven access to Amazon Web Services APIs, serving as the foundation for the AWS CLI and boto3 libraries.
Install it if you need programmatic access to AWS services.
Provides an async client for AWS services using botocore and aiohttp, allowing you to call AWS APIs asynchronously within asyncio-based applications.
Install it if you need to call AWS services from async Python code; it is the standard way to do so.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides a platform-independent file locking mechanism to coordinate access to files across processes and threads.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides common Protocol Buffer message definitions used across Google Cloud APIs, enabling Python clients to interact with Google services.
See also savepagenow · waybackpy · warc3-wet · openbb-crypto · rpmfile · aws-sso-util · warc3-wet-clueweb09 · pipx · yt-dlp · tarsafe