ultimate-sitemap-parser
A performant library for parsing and crawling sitemaps
Decision gist · record as of 2026-08-14
Yes. The library is actively maintained, has low install friction, no known vulnerabilities, and solves a specific problem well—parsing diverse sitemap formats reliably. The copyleft GPL-3.0-or-later license is the main constraint: use it freely in open-source projects, but review licensing implications before bundling into proprietary software.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; network access to fetch sitemaps from target URLs.
- Low friction: pure Python wheel with only two stable runtime dependencies (python-dateutil and requests).
- Actively maintained as of 2026-06-16 with 256 repository stars.
License · maintenance · safety
GPL-3.0-or-later (copyleft) — GPL-3.0-or-later (copyleft): you must license any derivative work or bundled application under compatible terms; suitable for open-source projects but requires legal review before use in proprietary software.
last release 2026-06-16 (59 days) · last repo commit 2026-06-16 · 256 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 198,401 downloads/mo, #9,731 on PyPI
Alternatives
Verify before relying
pip install ultimate-sitemap-parser
from ultimate_sitemap_parser.tree import sitemap_tree_for_homepage
tree = sitemap_tree_for_homepage('https://www.example.org/')
for page in tree.all_pages():
print(page.url)- Whether the library handles redirects or authentication when fetching sitemaps from protected URLs.
- Memory consumption profile on very large sitemap hierarchies beyond the ~1 million URLs mentioned in testing.
- Whether custom web client support extends to proxy configuration or certificate handling.
What it is and what it does
Ultimate Sitemap Parser is a Python library that discovers and parses sitemaps in all common formats—XML, RSS, Atom, plain text, and Google News/Image variants—and extracts URLs into an object tree. It handles malformed sitemaps gracefully, discovers sitemaps linked from robots.txt, and uses memory-efficient Expat XML parsing to avoid loading entire hierarchies into memory at once.
The library is designed for web crawlers and indexing workflows. You give it a homepage URL, and it returns a tree of sitemap objects you can iterate over to get all discovered pages. It has been field-tested with approximately 1 million URLs as part of the Media Cloud project and depends only on python-dateutil and requests, making installation straightforward on any modern Python environment.
Use it for
- Discover all pages on a website by parsing its sitemap hierarchy for web crawling or SEO audits.
- Extract URLs from Google News or Image sitemaps for specialized content indexing.
- Build a site map inventory by recursively following nested sitemap references.
- Integrate sitemap discovery into a web scraper to respect site structure and robots.txt directives.
- Analyze sitemap coverage to identify missing or orphaned pages in a website.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The library is actively maintained, has low install friction, no known vulnerabilities, and solves a specific problem well—parsing diverse sitemap formats reliably. The copyleft GPL-3.0-or-later license is the main constraint: use it freely in open-source projects, but review licensing implications before bundling into proprietary software.
Install
ultimate-sitemap-parser on PyPI
Before you install
Low friction: pure Python wheel with only two stable runtime dependencies (python-dateutil and requests). Actively maintained as of 2026-06-16 with 256 repository stars.
Requires Python 3.10 or later; network access to fetch sitemaps from target URLs.
License in practice
GPL-3.0-or-later (copyleft): you must license any derivative work or bundled application under compatible terms; suitable for open-source projects but requires legal review before use in proprietary software.
Quickstart
pip install ultimate-sitemap-parser
from ultimate_sitemap_parser.tree import sitemap_tree_for_homepage
tree = sitemap_tree_for_homepage('https://www.example.org/')
for page in tree.all_pages():
print(page.url)
Verify before relying
- Whether the library handles redirects or authentication when fetching sitemaps from protected URLs.
- Memory consumption profile on very large sitemap hierarchies beyond the ~1 million URLs mentioned in testing.
- Whether custom web client support extends to proxy configuration or certificate handling.
Package facts
| License | GPL-3.0-or-later copyleft |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespython-dateutilrequests |
| Maintenance | Actively maintained 59 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 198,401 / month, #9,731 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyLicense :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)Operating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTP :: Indexing/SearchTopic :: Text Processing :: IndexingTopic :: Text Processing :: Markup :: XML |
Evidence: ultimate_sitemap_parser-1.8.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “sitemap parser xml”
- ultimate-sitemap-parserParses and crawls sitemaps in multiple formats (XML, RSS, Atom, plain…
- sphinx-sitemapGenerates sitemaps.org-compliant XML sitemaps for Sphinx…
- advertoolsadvertools provides data manipulation and analysis functions for…
Give your agent the search over MCP, or paste the wish link into any chat.
More XML packages
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Defusedxml hardens Python's standard XML libraries against XML bomb attacks and entity expansion exploits by providing drop-in replacements that disable dangerous parsing features by default.
Install it if your application parses any XML from untrusted sources.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts XML to Python dictionaries and back, treating XML parsing and generation like working with JSON.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
See also sphinx-sitemap · googlenewsdecoder · feedparser · Protego · feedfinder2 · courlan · feedgen · readable-content · Crawl4AI · gnews