url-normalize
URL normalization for Python
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a well-defined problem. It is suitable for production use in URL deduplication, web crawling, and caching workflows where canonical URL representation matters.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low friction: single pure-Python dependency (idna), wheel distribution, and active maintenance with a recent release.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted commercial and private use with minimal obligations—suitable for most projects.
last release 2026-04-25 (111 days) · last repo commit 2026-04-25 · 100 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 22,421,877 downloads/mo, #972 on PyPI
Alternatives
Verify before relying
pip install url-normalize
from url_normalize import url_normalize
url_normalize("HTTP://User:Pass@www.FOO.com:80///foo/../bar/./baz?q=1#frag")
# 'http://User:Pass@www.foo.com/bar/baz?q=1#frag'- Whether query parameter filtering defaults match documented behavior across all edge cases.
- Performance characteristics on very large URL batches or deeply nested path segments.
- Stability of the IDNA2008 with UTS46 implementation across different Unicode versions.
- Default port handling for schemes other than http and https.
What it is and what it does
url-normalize is a URL canonicalization library that transforms URLs into a standardized form suitable for deduplication, caching, and comparison. It handles scheme and host lowercasing, removes default ports, resolves relative path segments (..), and applies RFC-compliant percent-encoding. It also supports internationalized domain names through idna, allowing proper normalization of non-ASCII domains.
The library is designed for scenarios where you need to treat semantically equivalent URLs as identical strings—database deduplication, web crawling, and link analysis are typical use cases. It offers configurable defaults (scheme, domain), query parameter filtering with allowlists, and a humanization function to convert normalized URLs back to user-readable form. A command-line interface is also available for standalone use.
Use it for
- Deduplicating URLs in a web crawl or link database by normalizing them to a canonical form before storage.
- Filtering tracking parameters from URLs while preserving functional query arguments via allowlists.
- Resolving relative URLs found on a page by supplying a default domain and scheme.
- Displaying normalized URLs in a human-readable format for IDN domains and percent-encoded paths.
- Comparing URLs for equivalence in caching or request deduplication without string comparison errors.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a well-defined problem. It is suitable for production use in URL deduplication, web crawling, and caching workflows where canonical URL representation matters.
Install
url-normalize on PyPI
Before you install
Low friction: single pure-Python dependency (idna), wheel distribution, and active maintenance with a recent release.
Requires Python 3.10 or later.
License in practice
MIT license permits unrestricted commercial and private use with minimal obligations—suitable for most projects.
Quickstart
pip install url-normalize
from url_normalize import url_normalize
url_normalize("HTTP://User:Pass@www.FOO.com:80///foo/../bar/./baz?q=1#frag")
# 'http://User:Pass@www.foo.com/bar/baz?q=1#frag'
Verify before relying
- Whether query parameter filtering defaults match documented behavior across all edge cases.
- Performance characteristics on very large URL batches or deeply nested path segments.
- Stability of the IDNA2008 with UTS46 implementation across different Unicode versions.
- Default port handling for schemes other than http and https.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packageidna |
| Maintenance | Actively maintained 111 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 22,421,877 / month, #972 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: url_normalize-3.0.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “canonical url”
- url-normalizeStandardizes URLs to a canonical form by lowercasing schemes and…
- urlpyurlpy parses, normalizes, and compares URLs with support for…
- s3urlsParse and build Amazon S3 URLs in multiple formats, extracting or…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also courlan · urlcanon · urlpy · furl · ada-url · w3lib · yarl · dsnparse · purl · multiurl