Subcategories
Packages
Parse, create, and manipulate JUnit/xUnit test result XML files, including merging multiple reports and supporting extended schemas like xunit2.
Install it if you need to parse, generate, or manipulate JUnit/xUnit XML test reports.
Computes Levenshtein edit distance, string similarity, and approximate median strings through a C extension module for fast string comparison operations.
Converts tabular data into formatted tables for output in multiple text, binary, and application-specific formats including Markdown, HTML, CSV, JSON, Excel, LaTeX, and SQLite.
Install it if you need to generate formatted tables in any of its supported output formats; skip it only if your use case is limited to a single format where a…
Represents and manipulates IPv4, IPv6, MAC addresses and related network objects; supports CIDR notation, subnetting, set operations, IANA lookups, and DNS reverse generation.
Parses Excel 2007-2010 Binary Workbook (xlsb) files and extracts data as rows of cells with values, dates, and positions.
However, do not rely on it for new projects or expect updates—it is abandoned.
Converts bidirectional text (mixed left-to-right and right-to-left scripts) into display order for rendering, with both Python and Rust implementations available.
Install it if you need to display bidirectional text; the LGPL license is standard for text-processing libraries and poses no barrier for most projects.
TheFuzz performs fuzzy string matching using Levenshtein Distance to find approximate matches between text sequences, with support for multiple matching strategies and batch processing.
However, note that it's aging—the last release was in January 2024—so verify whether the maintainers are still actively developing it or if it's in stable maintenance…
Parses HCL (HashiCorp Configuration Language) files into Python dictionaries, providing load/loads/dumps functions similar to the json module.
However, do not use it for modern Terraform (which requires HCL2 support); maintenance is aging with the last release in 2023-09-01, so expect limited support for new…
Parses PDF files into structured markdown or text format for use with LlamaIndex, enabling efficient document retrieval and context augmentation in RAG pipelines.
Implements BM25 ranking algorithms (Okapi BM25, BM25L, BM25+) to score and rank documents by relevance to a query.
Pypandoc wraps pandoc, a universal document converter, letting you convert between document formats (Markdown, HTML, DOCX, PDF, and many others) from Python code.
Install it if you need to convert between document formats from Python code.
Extracts text, headers, footers, hyperlinks, and images from Microsoft Word .docx files using pure Python, with both command-line and programmatic interfaces.
Python wrapper for wkhtmltopdf that converts HTML to PDF using the Webkit rendering engine, supporting URLs, files, and HTML strings as input.
Converts Rich Text Format (RTF) files to plain text, stripping formatting and control codes to produce readable output suitable for parsing and processing.
Install it if you need to process RTF files; skip it if your documents are already in plain text or modern formats.
Converts strings between different naming conventions (camelCase, snake_case, PascalCase, kebab-case, and others) with a simple function-per-case API.
However, the package is abandoned (last release 2017-08-06) and may fail on modern Python versions or edge cases.
Registers additional EBCDIC character encoding codecs for Python, enabling encode/decode operations with legacy mainframe systems that use EBCDIC character sets.
Decodes multi-byte character strings by detecting their encoding and converting them to Unicode, using chardet for automatic codec detection.
Install it if you need to reliably convert encoded bytes to Unicode without knowing the source encoding in advance.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
Python bindings for jq 1.8.2 that let you compile and execute jq programs against JSON data, with multiple input and output methods.
Compresses and decompresses Rich Text Format (RTF) data using the LZFu compression algorithm, implementing the Microsoft RTF Compression format.
However, maintenance is aging with no recent updates; verify that it meets your RTF variant requirements before relying on it for new projects.
Renders ASCII text as ASCII art using a variety of fonts, available both as a command-line tool and a Python library.
Install it if you need ASCII art text rendering in Python or at the command line; the low friction and stable API make it a straightforward addition to any project.
Parses incomplete JSON strings generated by LLM streaming, extracting usable data before the full response arrives, and can complete partial JSON to valid syntax.
Adds humanize library filters to Jinja2 templates, enabling human-readable formatting of numbers, dates, times, and file sizes directly in template expressions.
Drop-in replacement for Python 2.7's csv module that handles Unicode strings natively, eliminating encoding errors when reading and writing CSV data.
Reads Excel and ODF spreadsheet files via a Python binding to the Rust calamine library, returning sheet data as Python lists or compatible with pandas.
selectolax is a fast HTML5 parser with CSS selector support, written in Cython and backed by either the Modest or Lexbor C parsing engines.
Detects sentence boundaries in text using rule-based heuristics, splitting paragraphs into individual sentences while handling edge cases like abbreviations and decimal points.
However, do not install if you need active maintenance, multi-language support, or assurance of ongoing bug fixes.
Provides encoding and character detection utilities, building on chardet to handle text encoding operations.
However, it is brand-new (first release 2026-04-19) with minimal public documentation and only 1 star on GitHub.
Interegular checks whether pairs of Python regular expressions can match overlapping strings, converting regex patterns to finite-state machines to detect intersections.
Provides diff, match, and patch algorithms for comparing plain text, finding fuzzy matches, and applying patches—originally built for Google Docs and now maintained as a Python port of Google's Diff Match and Patch library.
Install it if you need robust diff, match, or patch operations for plain text—it's a well-established tool with low friction and proven algorithms.
Polyleven computes Levenshtein distance between two strings using a fast C implementation, with optional threshold support to skip expensive comparisons.
Extracts text, coordinates, and bitmap images from programmatic PDFs with support for character, word, and line-level granularity, offering both sequential and multi-threaded parsing modes.
Datefinder extracts date and time expressions from unstructured text and converts them into datetime objects or typed match objects with confidence and span information.
Provides file loaders that parse documents in multiple formats (PDF, DOCX, images, CSV, HTML, Markdown, and others) into structured data for indexing and retrieval.
Install it if you need to ingest documents from multiple file types.
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
Maps ISO 639 language codes (639-1, 639-2, 639-3) to language names and metadata, with methods to look up or guess a language from any code or name format.
SacreBLEU computes BLEU, chrF, and TER scores for machine translation evaluation with automatic test set management and reproducible, comparable results across systems.
Install it if you work with machine translation evaluation, benchmark scoring, or need comparable BLEU/chrF/TER metrics across systems.
Parses and executes JSONPath expressions against JSON data, returning matched values and their full paths within the document structure.
No, not for new projects.
Provides Python bindings to the Rust regress crate for ECMA-compliant regular expression matching, offering an alternative regex engine to Python's built-in re module.
However, the aging maintenance status and Beta classification mean you should verify that ECMA regex behavior is actually required for your use case—Python's built-in…
Provides Unicode normalization (NFC, NFD, NFKC, NFKD) using Unicode Standard version 17.0 data, independent of the host Python environment's built-in Unicode version.
Install it if your application requires consistent normalization across different systems or if you need to target a specific Unicode Standard version.