Packages
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Defusedxml hardens Python's standard XML libraries against XML bomb attacks and entity expansion exploits by providing drop-in replacements that disable dangerous parsing features by default.
Install it if your application parses any XML from untrusted sources.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts XML to Python dictionaries and back, treating XML parsing and generation like working with JSON.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Parses Atom and RSS feeds (including RSS 0.9x, RSS 1.0, RSS 2.0, CDF, Atom 0.3, and Atom 1.0) into Python data structures.
Install it if you need to consume RSS or Atom feeds.
Trafilatura extracts main text, metadata, and structured content from web pages and HTML, converting raw HTML into clean, usable data in multiple output formats.
Python bindings for XML Security Library, enabling cryptographic signing and verification of XML documents.
Provides XPath 1.0, 2.0, 3.0, and 3.1 selectors for querying ElementTree and lxml XML data structures using standard XPath expressions.
Install it if you need to query XML with XPath expressions beyond version 1.0.
Validates XML documents against XSD schemas and converts XML to/from Python data structures and JSON.
Python wrapper for wkhtmltopdf that converts HTML to PDF using the Webkit rendering engine, supporting URLs, files, and HTML strings as input.
Svglib reads SVG files and converts them to ReportLab Drawing objects, which can then be rendered to PDF, bitmap, or other formats supported by ReportLab.
Parsel extracts data from HTML, JSON, and XML documents using CSS selectors, XPath expressions, JMESPath queries, and regular expressions.
Install it if you need to extract data from HTML, XML, or JSON documents in a Python application.
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
xsdata generates Python dataclasses from XML schemas, WSDL, DTD, and JSON documents, then parses and serializes XML and JSON data using those generated models.
Pydantic-xml adds XML serialization and deserialization to pydantic models, letting you bind model fields to XML attributes, elements, and text content, then convert between Python objects and XML documents.
DocLang is a reference toolkit for validating and packaging documents in DocLang format, an AI-native markup designed to preserve structure and semantics while mapping cleanly to LLM tokens.
Pybtex reads bibliography data from BibTeX, BibTeXML, or YAML files and generates formatted bibliographies in LaTeX, HTML, markdown, or plain text, supporting both BibTeX style files and Python-based custom styles.
Install it if you need to process BibTeX files programmatically, generate bibliographies in multiple formats, or integrate bibliography handling into a Python workflow.
A dict subclass that adds dot-notation access, keypath queries, and normalized I/O for formats like JSON, YAML, CSV, XML, and others.
Install it if you work frequently with nested dicts, multi-format data loading, or want cleaner syntax for accessing deeply nested values—the keypath and I/O features…
Converts Python dictionaries into XML strings with configurable formatting, wrapping, and tag behavior.
Install it if you need straightforward dictionary-to-XML conversion.
Extends Python's unittest framework to compare XML documents structurally using XPath expressions and lxml, rather than comparing raw XML strings.
Integrates BibTeX citations into docutils-generated documentation, enabling bibliography rendering in reStructuredText documents without Sphinx.
However, maintenance is aging (last release 1088 days ago); verify compatibility with your docutils version before relying on it in production.
MarkupPy generates HTML and XML markup from Python code using an intuitive, pythonic API without external dependencies.
Install it if you need to build markup from Python code without external dependencies or template syntax overhead.
Reformats HTML and XML strings with intelligent inline-tag handling, avoiding the excessive line breaks that tools like BeautifulSoup.prettify() introduce between tags that should stay on the same line.
imgkit wraps the wkhtmltoimage command-line tool to convert HTML (from URLs, files, or strings) into image files using the Webkit rendering engine.
However, do not use it in production without understanding that no upstream maintenance is available—if wkhtmltoimage itself breaks or you encounter bugs in imgkit,…
Converts XML documents to JSON format via a Python library or command-line tool, using xmltodict as its underlying parser.
Junos PyEZ is a Python library for remotely managing and automating Junos devices via NETCONF, enabling both interactive scripting and programmatic control without requiring deep Junos XML API knowledge.
Install it if you need to programmatically manage Junos devices; skip it if you have no Junos devices or prefer CLI-only management.
xmldiff compares two XML files or trees and produces a human-readable diff that accounts for the hierarchical structure of XML, rather than treating it as flat text.
Automate programmatic interaction with HTTP web servers by simulating a stateful browser—fill forms, follow links, manage history, and parse HTML without a GUI.
However, consider that maintenance is aging—last release was 840 days ago—so evaluate whether the package meets your Python version and modern web compatibility needs…
PyXB-X generates Python classes from XMLSchema definitions, enabling bidirectional conversion between XML documents and Python objects with a Pythonic interface.
Generates web feeds in ATOM and RSS formats, with support for extensions including podcast feeds.
Generates pydantic models from XML and JSON schemas, then parses and serializes XML/JSON documents as typed Python objects using those models.
Arelle is an end-to-end XBRL processor that validates, parses, and transforms financial reporting documents in XBRL format, accessible via GUI, CLI, Python API, and web service.
Install it if you need to validate, parse, or transform XBRL financial reports—especially for SEC, EU, or other regulatory filings.
Genshi is a Python library for parsing, generating, and processing HTML, XML, and other textual content, with a built-in template language inspired by Kid for web output generation.
However, verify Python version support for your target runtime, as classifiers list both Python 2 and 3 without specifying which versions are actually supported.
Reads, writes, and manipulates KML files—an XML-based geospatial data format—using a simple Python API with support for geometry objects and date/time handling.
Converts XML to Python data structures (lists, dicts, strings) that preserve XML metadata, and converts Python objects back to XML.
However, do not use it for new projects requiring active maintenance, security updates, or support for modern Python versions beyond 3.5.
Parses and crawls sitemaps in multiple formats (XML, RSS, Atom, plain text, Google News/Image) and extracts URLs efficiently without loading entire trees into memory.
Sickle is a lightweight Python client library for harvesting metadata from OAI-PMH (Open Archives Initiative Protocol for Metadata Harvesting) compliant repositories, handling all six OAI verbs and automatically deserializing Dublin Core metadata.
However, verify that your target repositories still use OAI-PMH and that you are comfortable with no future updates or support for newer Python versions.
Reads and writes biomedical text annotation formats—BioC XML/JSON, Brat standoff, and PubTator—with a marshal/pickle-like API.
Install only if you are confident the format specifications and your dependencies will remain compatible, or if you plan to maintain a fork yourself.