Packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Converts numbers, dates, times, and file sizes into human-readable text formats, with support for fuzzy durations like "3 minutes ago" and localization to multiple languages.
Install it if you need to display human-readable numbers, durations, or sizes to end users.
Formats and parses numbers, file sizes, timespans, and other values into human-readable text; provides terminal interaction utilities including ANSI text styling and user prompts.
Python wrapper for wkhtmltopdf that converts HTML to PDF using the Webkit rendering engine, supporting URLs, files, and HTML strings as input.
Parsimonious is a pure-Python PEG (parsing expression grammar) parser that builds abstract syntax trees from text according to grammar rules you define, with no external lexer or parser generator required.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Parses, validates, and reformats standard numbers and codes across many countries and industries—tax IDs, bank accounts, identity numbers, VAT codes, and financial identifiers.
Install it if you validate or reformat standardized numbers.
Extracts tables from PDF files and converts them into pandas DataFrames, CSV, TSV, or JSON formats using a Java-based backend.
However, dormant development since 2024-10-17 means new features or breaking-change fixes are unlikely.
Converts between LaTeX code and Unicode text in both directions, parsing LaTeX markup into a logical structure and encoding Unicode strings into LaTeX escape sequences.
Install it if you work with LaTeX text processing, bibliographic data, or need bidirectional LaTeX-Unicode conversion.
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
Maps ISO 639 language codes (639-1, 639-2, 639-3) to language names and metadata, with methods to look up or guess a language from any code or name format.
Converts Unicode text to ASCII-only equivalents by transliterating each character to its best ASCII representation, leaving ASCII input unchanged and removing unmapped characters.
semchunk splits text into semantically meaningful chunks while preserving local context, supporting custom tokenizers, chunk overlapping, offsets, and optional AI-powered chunking via the Isaacus API.
Install it if you need to chunk text for RAG, embeddings, or language model workflows.
Lark is a parsing library that builds abstract syntax trees from context-free grammars using Earley, LALR(1), or CYK parsing algorithms, with minimal code required.
jaconv converts between Japanese character types: Hiragana, Katakana, half-width (Hankaku) and full-width (Zenkaku) characters, plus transliteration to and from romanized alphabet.
Install it if you work with Japanese text preprocessing, normalization, or transliteration.
StringZilla provides SIMD and SWAR-accelerated string operations including substring search, hashing, edit distances, sorting, and segmentation for Python, with no runtime dependencies.
Install only if your Python version is 3.10 or later.
Parses and serializes JSON5 (a more lenient JSON variant) and standard JSON in Python, with output safe for HTML templates.
Install it if you need JSON5 support or HTML-safe serialization; otherwise the standard json module suffices.
Converts text and generates one-line ASCII art symbols, rendering them as multi-line ASCII text using various fonts and decorative styles.
PyThaiNLP provides Thai-language natural language processing tools including tokenization, part-of-speech tagging, transliteration, spelling correction, and linguistic utilities, designed as a Thai counterpart to NLTK.
Install it if you need to process Thai text; the base package is lightweight and the optional extras allow you to add machine translation or WordNet support as needed.
Provides Python bindings to TREC's trec_eval tool for computing standard Information Retrieval evaluation measures like MAP and NDCG on ranked search results.
Computes Jaro and Jaro-Winkler string similarity scores, returning values from 0 (no match) to 1 (perfect match) for comparing two strings.
Compresses CSS by removing unnecessary whitespace and comments, following the YUI CSS Compressor algorithm.
TatSu compiles extended EBNF grammars into memoizing PEG parsers in Python, or generates Python source code that implements those parsers.
Install it if you need to parse hierarchical data, build a DSL, or generate parsers from grammars.
EmPy is a templating system that embeds Python code directly into text documents using customizable markup (default `@`), processing the combined source to produce output with Python expressions, statements, and control structures evaluated inline.
Converts text in any script to Latin alphabet romanization, supporting multiple languages with context-aware character mappings and optional language codes.
However, the dormant maintenance status (no activity for 777 days) means you should verify that its romanization rules and Unicode support remain adequate for your…
imgkit wraps the wkhtmltoimage command-line tool to convert HTML (from URLs, files, or strings) into image files using the Webkit rendering engine.
However, do not use it in production without understanding that no upstream maintenance is available—if wkhtmltoimage itself breaks or you encounter bugs in imgkit,…
Esprima parses ECMAScript (JavaScript) source code into an abstract syntax tree, supporting tokenization and full syntactic analysis of JavaScript programs.
Install only if you are maintaining legacy code that already depends on esprima and cannot migrate.
Provides a Django model field that automatically generates and maintains URL-friendly slugs from other fields, with built-in uniqueness enforcement and support for custom slugification logic.
However, the last release was April 2023 and maintenance is aging; verify compatibility with your current Django version before adopting.
Provides detailed Unicode character properties and metadata from the Unicode Character Database with human-readable aliases, as an alternative to Python's standard library unicodedata module.
No, not recommended for new projects.
A Sphinx theme that extends PyData Sphinx Theme with NVIDIA branding and styling for documentation projects.
However, the proprietary license is restrictive—verify that your project qualifies as a 'NVIDIA product or service' before committing.
Compares two JSON files and generates a JSON output showing the differences, with options to include or exclude specific keys from the comparison.
Install only if you have a specific legacy requirement and can accept the lack of support.
Computes phonetic keys of strings using Soundex, NYSIIS, Metaphone, and Double Metaphone algorithms for fuzzy matching and phonetic indexing.
Detects whether strings are gibberish or valid text by training on example corpora and scoring input against learned patterns.
Parses and tokenizes JavaScript code up to ECMA 2025 syntax, producing an ESTree-compatible abstract syntax tree or token stream for analysis and manipulation.
Efficiently parse and stream-process MediaWiki XML database dumps with memory-conscious iteration and optional distributed processing across multiple files.
Install it if you need to process wiki database exports; skip it if you're not working with MediaWiki XML data.
Provides standardized Python classes for MediaWiki data types, normalizing representations across database, XML dumps, and API responses.
However, verify that the included types match your specific use case, and be aware that maintenance is dormant—if you need active support or compatibility updates for…