$npx skillfedfor your agent

urlextract

Collects and extracts URLs from given text.

With conditionsPyPI Python ModulesReleased Feb 20241.2M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — urlextract-1.9.0-py3-none-any.whl
v1.9.0 · released 2024-02-29 · 4 runtime deps: idna, uritools, platformdirs, filelock

Yes, if you need straightforward URL extraction from unstructured text. The package is stable (Production/Stable status), has no known security vulnerabilities, and carries a permissive MIT license. Install friction is low. However, the last release was 897 days ago and maintenance is dormant—if you need active support or expect frequent TLD updates, verify that the cached list meets your requirements or plan to fork/maintain it yourself.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low friction install with four lightweight runtime dependencies.
  • Last release was 897 days ago; repository is not archived but marked dormant, indicating the package is stable but receives infrequent updates.

License · maintenance · safety

MIT (permissive) — MIT license is permissive—you can use, modify, and distribute this package freely in commercial and open-source projects with minimal restrictions.

last release 2024-02-29 (897 days) · last repo commit 2024-02-29 · 278 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,153,061 downloads/mo, #4,293 on PyPI

Verify before relying

pip install urlextract

from urlextract import URLExtract

extractor = URLExtract()
urls = extractor.find_urls("Text with URLs. Let's have URL janlipovsky.cz as an example.")
print(urls)  # prints: ['janlipovsky.cz']
  • Whether DNS validation via dnspython is included in the base install or requires separate setup
  • Current accuracy of TLD list and how often it is refreshed from iana.org
  • Performance characteristics on very large texts or high-volume extraction scenarios
Same gist for agents: .md · .json

What it is and what it does

URLExtract is a Python library that finds URLs embedded in plain text by scanning for valid top-level domains (TLDs) and expanding outward to locate word boundaries. It works by locating any TLD occurrence in the input text, then searching left and right from that position until it hits a stop character like whitespace or punctuation. The library maintains an up-to-date TLD list downloaded from iana.org and can optionally validate extracted domains via DNS checks using uritools and idna for proper domain name handling.

The package provides three main extraction modes: `find_urls()` returns a list of all URLs found, `gen_urls()` yields URLs as a generator for memory efficiency, and `has_urls()` performs a quick boolean check. You can also manually update the TLD cache or set it to auto-update after a specified number of days. The library is aware of a known limitation: since some TLDs are also valid English words, it may produce false positives in contexts like CSS class selectors (e.g., detecting 'p.bold.name' as a domain).

Use it for

  • Extract clickable links from user-generated text, chat messages, or social media posts for processing or display
  • Scan HTML or text documents to identify and collect all URLs for link validation or archival purposes
  • Pre-process raw text data before feeding it to NLP or machine learning pipelines that need URL detection
  • Build web scrapers or crawlers that need to identify target URLs within fetched page content
  • Validate or sanitize user input by detecting whether a text block contains any URLs

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need straightforward URL extraction from unstructured text.

The package is stable (Production/Stable status), has no known security vulnerabilities, and carries a permissive MIT license. Install friction is low. However, the last release was 897 days ago and maintenance is dormant—if you need active support or expect frequent TLD updates, verify that the cached list meets your requirements or plan to fork/maintain it yourself.

Install

urlextract on PyPI

Before you install

Low friction install with four lightweight runtime dependencies. Last release was 897 days ago; repository is not archived but marked dormant, indicating the package is stable but receives infrequent updates.

License in practice

MIT license is permissive—you can use, modify, and distribute this package freely in commercial and open-source projects with minimal restrictions.

Quickstart

pip install urlextract

from urlextract import URLExtract

extractor = URLExtract()
urls = extractor.find_urls("Text with URLs. Let's have URL janlipovsky.cz as an example.")
print(urls)  # prints: ['janlipovsky.cz']

Verify before relying

  • Whether DNS validation via dnspython is included in the base install or requires separate setup
  • Current accuracy of TLD list and how often it is refreshed from iana.org
  • Performance characteristics on very large texts or high-volume extraction scenarios

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
idnauritoolsplatformdirsfilelock
MaintenanceDormant 897 days since the last release
Last repo commit
First released
Downloads1,153,061 / month, #4,293 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Software Development :: Libraries :: Python ModulesTopic :: Text ProcessingTopic :: Text Processing :: Markup :: HTML

Evidence: urlextract-1.9.0-py3-none-any.whl

Tags

Capabilities
extract urls from textfind links in texturl detectiontld-based url extractionparse urls from stringscollect urls from contenturl finder
Topics
url-extractiontext-processingtld-parsing
PyPI keywords
urlextractfindfindercollectlinktldlist

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract urls from text”

  • urlextractExtracts URLs from text by locating TLDs and expanding boundaries to…
  • commonregexExtracts dates, times, emails, phone numbers, links, IP addresses,…
  • micawberExtracts rich metadata (title, author, thumbnail, embed HTML) from…

Give your agent the search over MCP, or paste the wish link into any chat.

More Python Modules packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
PyYAML Worth it
PyPI · Python Modules · released Sep 2025

PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.

MITcompiled wheel · 3.8+
1.2Bdownloads / mo
pydantic Worth it
PyPI · Python Modules · released May 2026

Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.

MITpure Python · 3.9+
1.1Bdownloads / mo
annotated-types Worth it
PyPI · Python Modules · released Jul 2026

Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.

Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…

MITpure Python · 3.10+
871.3Mdownloads / mo
typing-inspection Worth it
PyPI · Python Modules · released Aug 2026

Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.

MITpure Python · 3.10+
783.0Mdownloads / mo

See also mozilla-repo-urls · tldextract · tld · tldparse · tlds · linkify-it-py · publicsuffix2 · query-string · textract · ldapdomaindump