{"categories":[{"label":"WWW/HTTP","url":"https://skillfed.io/packages/category/internet-www-http/3"}],"enrichment":{"capability":"Resiliparse provides fast, robust parsing and analysis tools for web archive data, written in Rust and C++ with Python bindings, and includes support for processing web content through its fastwarc dependency.","skillfed_tags":["web-archive","warc-processing","rust-bindings"],"use_cases":["Parse and extract content from WARC (Web ARChive) files at scale for research or analytics.","Process Common Crawl data or other large web archives for information extraction pipelines.","Build web analytics and data mining workflows that require efficient HTML and metadata parsing.","Integrate web archive processing into data science or NLP projects that need robust, fast parsing."],"what_it_does":"Resiliparse is a collection of parsing and analysis tools for web archive data, built with Rust and C++ for performance and wrapped with Python bindings. It is part of the ChatNoir web analytics toolkit and depends on fastwarc for WARC file handling. The package is designed to handle large-scale web archive processing efficiently, with pre-built binaries available for supported platforms.\n\nThe library targets developers and researchers working with web crawl data, Common Crawl archives, and similar large-scale web datasets. Installation is straightforward via pip for supported platforms; building from source requires vcpkg and C++ tooling but is documented. The project is actively maintained and carries no known security vulnerabilities.","worth_installing":"Yes, if you work with web archives or WARC files. Resiliparse is actively maintained, carries no known vulnerabilities, and offers pre-built wheels for supported platforms. Medium install friction is acceptable for the performance and robustness it provides. The permissive Apache-2.0 license poses no restrictions. Not relevant if you don't need web archive processing."},"id":"resiliparse","links":{"html":"https://skillfed.io/packages/resiliparse","md":"https://skillfed.io/packages/resiliparse.md","pypi":"https://pypi.org/project/resiliparse/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-07-20","license_spdx":"Apache-2.0","license_treatment":"permissive","name":"Resiliparse","python_support":"supports_current","summary":"A collection of robust and fast processing tools for parsing and analyzing (not only) web archive data."},"popularity":{"monthly_downloads":1317625,"position":4067,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"1.0.9"}
