$npx skillfedfor your agent

ultimate-sitemap-parser

A performant library for parsing and crawling sitemaps

Worth itPyPI XMLReleased Jun 2026198.4K downloads / moGPL-3.0-or-laterPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — ultimate_sitemap_parser-1.8.1-py3-none-any.whl
v1.8.1 · released 2026-06-16 · Python >=3.10 · 2 runtime deps: python-dateutil, requests

Yes. The library is actively maintained, has low install friction, no known vulnerabilities, and solves a specific problem well—parsing diverse sitemap formats reliably. The copyleft GPL-3.0-or-later license is the main constraint: use it freely in open-source projects, but review licensing implications before bundling into proprietary software.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; network access to fetch sitemaps from target URLs.
  • Low friction: pure Python wheel with only two stable runtime dependencies (python-dateutil and requests).
  • Actively maintained as of 2026-06-16 with 256 repository stars.

License · maintenance · safety

GPL-3.0-or-later (copyleft) — GPL-3.0-or-later (copyleft): you must license any derivative work or bundled application under compatible terms; suitable for open-source projects but requires legal review before use in proprietary software.

last release 2026-06-16 (59 days) · last repo commit 2026-06-16 · 256 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 198,401 downloads/mo, #9,731 on PyPI

Verify before relying

pip install ultimate-sitemap-parser

from ultimate_sitemap_parser.tree import sitemap_tree_for_homepage

tree = sitemap_tree_for_homepage('https://www.example.org/')
for page in tree.all_pages():
    print(page.url)
  • Whether the library handles redirects or authentication when fetching sitemaps from protected URLs.
  • Memory consumption profile on very large sitemap hierarchies beyond the ~1 million URLs mentioned in testing.
  • Whether custom web client support extends to proxy configuration or certificate handling.
Same gist for agents: .md · .json

What it is and what it does

Ultimate Sitemap Parser is a Python library that discovers and parses sitemaps in all common formats—XML, RSS, Atom, plain text, and Google News/Image variants—and extracts URLs into an object tree. It handles malformed sitemaps gracefully, discovers sitemaps linked from robots.txt, and uses memory-efficient Expat XML parsing to avoid loading entire hierarchies into memory at once.

The library is designed for web crawlers and indexing workflows. You give it a homepage URL, and it returns a tree of sitemap objects you can iterate over to get all discovered pages. It has been field-tested with approximately 1 million URLs as part of the Media Cloud project and depends only on python-dateutil and requests, making installation straightforward on any modern Python environment.

Use it for

  • Discover all pages on a website by parsing its sitemap hierarchy for web crawling or SEO audits.
  • Extract URLs from Google News or Image sitemaps for specialized content indexing.
  • Build a site map inventory by recursively following nested sitemap references.
  • Integrate sitemap discovery into a web scraper to respect site structure and robots.txt directives.
  • Analyze sitemap coverage to identify missing or orphaned pages in a website.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The library is actively maintained, has low install friction, no known vulnerabilities, and solves a specific problem well—parsing diverse sitemap formats reliably. The copyleft GPL-3.0-or-later license is the main constraint: use it freely in open-source projects, but review licensing implications before bundling into proprietary software.

Install

ultimate-sitemap-parser on PyPI

Before you install

Low friction: pure Python wheel with only two stable runtime dependencies (python-dateutil and requests). Actively maintained as of 2026-06-16 with 256 repository stars.

Requires Python 3.10 or later; network access to fetch sitemaps from target URLs.

License in practice

GPL-3.0-or-later (copyleft): you must license any derivative work or bundled application under compatible terms; suitable for open-source projects but requires legal review before use in proprietary software.

Quickstart

pip install ultimate-sitemap-parser

from ultimate_sitemap_parser.tree import sitemap_tree_for_homepage

tree = sitemap_tree_for_homepage('https://www.example.org/')
for page in tree.all_pages():
    print(page.url)

Verify before relying

  • Whether the library handles redirects or authentication when fetching sitemaps from protected URLs.
  • Memory consumption profile on very large sitemap hierarchies beyond the ~1 million URLs mentioned in testing.
  • Whether custom web client support extends to proxy configuration or certificate handling.

Package facts

LicenseGPL-3.0-or-later copyleft
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
python-dateutilrequests
MaintenanceActively maintained 59 days since the last release
Last repo commit
First released
Downloads198,401 / month, #9,731 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyLicense :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)Operating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTP :: Indexing/SearchTopic :: Text Processing :: IndexingTopic :: Text Processing :: Markup :: XML

Evidence: ultimate_sitemap_parser-1.8.1-py3-none-any.whl

Tags

Capabilities
sitemap parser xmlcrawl sitemaps pythonextract urls from sitemapsitemap tree parsinggoogle news sitemap parserrss atom sitemaprobots.txt sitemap discovery
Topics
web-crawlingsitemap-discoveryxml-parsing
PyPI keywords
sitemapcrawlerindexingxmlrssatomgoogle news

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “sitemap parser xml”

  • ultimate-sitemap-parserParses and crawls sitemaps in multiple formats (XML, RSS, Atom, plain…
  • sphinx-sitemapGenerates sitemaps.org-compliant XML sitemaps for Sphinx…
  • advertoolsadvertools provides data manipulation and analysis functions for…

Give your agent the search over MCP, or paste the wish link into any chat.

More XML packages

beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
defusedxml Worth it
PyPI · XML · released Mar 2021

Defusedxml hardens Python's standard XML libraries against XML bomb attacks and entity expansion exploits by providing drop-in replacements that disable dangerous parsing features by default.

Install it if your application parses any XML from untrusted sources.

permissive licensepure Pythondormant
241.2Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
xmltodict Worth it
PyPI · XML · released Feb 2026

Converts XML to Python dictionaries and back, treating XML parsing and generation like working with JSON.

MITpure Python · 3.9+
127.1Mdownloads / mo
Sphinx Worth it
PyPI · Software Development · released Dec 2025

Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.

BSD-2-Clausepure Python · 3.12+
91.8Mdownloads / mo

See also sphinx-sitemap · googlenewsdecoder · feedparser · Protego · feedfinder2 · courlan · feedgen · readable-content · Crawl4AI · gnews