courlan
Clean, filter and sample URLs to optimize data collection – includes spam, content type and language filters.
Decision gist · record as of 2026-08-14
Yes. Courlan is actively maintained, has no known vulnerabilities, and offers a focused, well-scoped toolkit for URL handling in crawling workflows. The permissive Apache-2.0 license and low install friction make it a practical choice for web scraping and document collection projects. Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or higher; earlier versions (3.6–3.9) require older courlan releases.
- Low install friction with three lightweight runtime dependencies (babel, tld, urllib3).
- Active maintenance with a recent release 74 days ago and ongoing repository activity.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most projects.
last release 2026-06-01 (74 days) · last repo commit 2026-08-12 · 177 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 13,207,347 downloads/mo, #1,292 on PyPI
Alternatives
Verify before relying
pip install courlan
from courlan import check_url, sample_urls, extract_links
# Validate and clean a URL
url, domain = check_url('https://example.org/page?utm_source=twitter')
# Sample URLs from a collection
sampled = sample_urls(['https://example.org/' + str(i) for i in range(100)], 10)
# Extract and filter links from HTML
links, priority_links = extract_links('<a href="/page">link</a>', 'https://example.org')- Performance characteristics when processing large URL collections (millions of URLs).
- Accuracy of language detection and filtering heuristics across different URL patterns.
- Whether the package handles internationalized domain names (IDN) correctly.
What it is and what it does
Courlan is a URL processing library designed to improve web crawler efficiency and document collection quality. It provides validation, normalization, filtering, and sampling functions that work together to identify high-value pages and discard low-quality content before crawling begins. The library filters out spam, tracking parameters, and platform-specific pages, and can apply language-aware heuristics to target specific locales.
The package is built on three runtime dependencies (babel, tld, urllib3) and offers both programmatic Python functions and command-line utilities. It includes specialized tools for crawl frontier management—distinguishing navigation pages from content pages, detecting non-crawlable URLs, and handling relative link resolution—making it useful for anyone building a web crawler, scraper, or bulk document collection pipeline.
Use it for
- Filter out tracker parameters and spam domains before crawling to reduce bandwidth waste and improve document quality.
- Sample representative URLs from a large domain to avoid redundant crawling and focus on diverse content.
- Extract and prioritize links from HTML with language filtering to target content in specific locales.
- Normalize and deduplicate URLs across multiple sources to prevent re-crawling the same page.
- Identify navigation and overview pages separately from content pages to optimize crawl scheduling.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Courlan is actively maintained, has no known vulnerabilities, and offers a focused, well-scoped toolkit for URL handling in crawling workflows. The permissive Apache-2.0 license and low install friction make it a practical choice for web scraping and document collection projects. Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.
Install
courlan on PyPI
Before you install
Low install friction with three lightweight runtime dependencies (babel, tld, urllib3). Active maintenance with a recent release 74 days ago and ongoing repository activity.
Requires Python 3.10 or higher; earlier versions (3.6–3.9) require older courlan releases.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most projects.
Quickstart
pip install courlan
from courlan import check_url, sample_urls, extract_links
# Validate and clean a URL
url, domain = check_url('https://example.org/page?utm_source=twitter')
# Sample URLs from a collection
sampled = sample_urls(['https://example.org/' + str(i) for i in range(100)], 10)
# Extract and filter links from HTML
links, priority_links = extract_links('<a href="/page">link</a>', 'https://example.org')
Verify before relying
- Performance characteristics when processing large URL collections (millions of URLs).
- Accuracy of language detection and filtering heuristics across different URL patterns.
- Whether the package handles internationalized domain names (IDN) correctly.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesbabeltldurllib3 |
| Maintenance | Actively maintained 74 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 13,207,347 / month, #1,292 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTPTopic :: Scientific/Engineering :: Information AnalysisTopic :: SecurityTopic :: Text Processing :: FiltersTopic :: Text Processing :: LinguisticTyping :: Typed |
Evidence: courlan-1.4.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “url validation and normalization”
- courlanValidates, normalizes, filters, and samples URLs for web crawling and…
- URLObjectURLObject provides a utility class for parsing, constructing, and…
- ada-urlParse, validate, and manipulate URLs according to the WHATWG URL…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also LinkChecker · url-normalize · Crawl4AI · trafilatura · ada-url · ultimate-sitemap-parser · crawlee · urlextract · scrapling · w3lib