$npx skillfedfor your agent

courlan

Clean, filter and sample URLs to optimize data collection – includes spam, content type and language filters.

Worth itPyPI WWW/HTTPReleased Jun 202613.2M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — courlan-1.4.0-py3-none-any.whl
v1.4.0 · released 2026-06-01 · Python >=3.10 · 3 runtime deps: babel, tld, urllib3

Yes. Courlan is actively maintained, has no known vulnerabilities, and offers a focused, well-scoped toolkit for URL handling in crawling workflows. The permissive Apache-2.0 license and low install friction make it a practical choice for web scraping and document collection projects. Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or higher; earlier versions (3.6–3.9) require older courlan releases.
  • Low install friction with three lightweight runtime dependencies (babel, tld, urllib3).
  • Active maintenance with a recent release 74 days ago and ongoing repository activity.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most projects.

last release 2026-06-01 (74 days) · last repo commit 2026-08-12 · 177 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 13,207,347 downloads/mo, #1,292 on PyPI

Verify before relying

pip install courlan

from courlan import check_url, sample_urls, extract_links

# Validate and clean a URL
url, domain = check_url('https://example.org/page?utm_source=twitter')

# Sample URLs from a collection
sampled = sample_urls(['https://example.org/' + str(i) for i in range(100)], 10)

# Extract and filter links from HTML
links, priority_links = extract_links('<a href="/page">link</a>', 'https://example.org')
  • Performance characteristics when processing large URL collections (millions of URLs).
  • Accuracy of language detection and filtering heuristics across different URL patterns.
  • Whether the package handles internationalized domain names (IDN) correctly.
Same gist for agents: .md · .json

What it is and what it does

Courlan is a URL processing library designed to improve web crawler efficiency and document collection quality. It provides validation, normalization, filtering, and sampling functions that work together to identify high-value pages and discard low-quality content before crawling begins. The library filters out spam, tracking parameters, and platform-specific pages, and can apply language-aware heuristics to target specific locales.

The package is built on three runtime dependencies (babel, tld, urllib3) and offers both programmatic Python functions and command-line utilities. It includes specialized tools for crawl frontier management—distinguishing navigation pages from content pages, detecting non-crawlable URLs, and handling relative link resolution—making it useful for anyone building a web crawler, scraper, or bulk document collection pipeline.

Use it for

  • Filter out tracker parameters and spam domains before crawling to reduce bandwidth waste and improve document quality.
  • Sample representative URLs from a large domain to avoid redundant crawling and focus on diverse content.
  • Extract and prioritize links from HTML with language filtering to target content in specific locales.
  • Normalize and deduplicate URLs across multiple sources to prevent re-crawling the same page.
  • Identify navigation and overview pages separately from content pages to optimize crawl scheduling.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Courlan is actively maintained, has no known vulnerabilities, and offers a focused, well-scoped toolkit for URL handling in crawling workflows. The permissive Apache-2.0 license and low install friction make it a practical choice for web scraping and document collection projects. Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.

Install

courlan on PyPI

Before you install

Low install friction with three lightweight runtime dependencies (babel, tld, urllib3). Active maintenance with a recent release 74 days ago and ongoing repository activity.

Requires Python 3.10 or higher; earlier versions (3.6–3.9) require older courlan releases.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; suitable for most projects.

Quickstart

pip install courlan

from courlan import check_url, sample_urls, extract_links

# Validate and clean a URL
url, domain = check_url('https://example.org/page?utm_source=twitter')

# Sample URLs from a collection
sampled = sample_urls(['https://example.org/' + str(i) for i in range(100)], 10)

# Extract and filter links from HTML
links, priority_links = extract_links('<a href="/page">link</a>', 'https://example.org')

Verify before relying

  • Performance characteristics when processing large URL collections (millions of URLs).
  • Accuracy of language detection and filtering heuristics across different URL patterns.
  • Whether the package handles internationalized domain names (IDN) correctly.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
babeltldurllib3
MaintenanceActively maintained 74 days since the last release
Last repo commit
First released
Downloads13,207,347 / month, #1,292 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTPTopic :: Scientific/Engineering :: Information AnalysisTopic :: SecurityTopic :: Text Processing :: FiltersTopic :: Text Processing :: LinguisticTyping :: Typed

Evidence: courlan-1.4.0-py3-none-any.whl

Tags

Capabilities
url validation and normalizationweb crawler url filteringspam and tracker removallanguage-aware url filteringurl deduplication and samplingcrawl frontier managementurl canonicalization
Topics
web-crawlingurl-processingdata-quality
PyPI keywords
cleanercrawleruriurl-parsingurl-manipulationurlsvalidationwebcrawling

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “url validation and normalization”

  • courlanValidates, normalizes, filters, and samples URLs for web crawling and…
  • URLObjectURLObject provides a utility class for parsing, constructing, and…
  • ada-urlParse, validate, and manipulate URLs according to the WHATWG URL…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also LinkChecker · url-normalize · Crawl4AI · trafilatura · ada-url · ultimate-sitemap-parser · crawlee · urlextract · scrapling · w3lib