Subcategories
Packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
Provides Python bindings to the Brotli compression library, enabling lossless compression and decompression using a modern LZ77 variant with Huffman coding and context modeling.
Install it if you need Brotli compression in Python; the API is straightforward and the library is battle-tested.
LlamaParse parses complex documents (PDFs, Word, Excel, PowerPoint, HTML) into structured text, markdown, or JSON optimized for RAG and LLM workflows, with built-in support for tables, images, and custom parsing instructions.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Converts text strings—including unicode, HTML entities, and special characters—into URL-safe slugs with customizable separators, length limits, and word boundaries.
Provides Pygments syntax highlighters for IPython code, console sessions, and special syntax like magics and shell commands.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Extracts text, images, and layout information from PDF documents by parsing the PDF source code directly, supporting modern PDF specifications and various encodings.
Converts numbers, dates, times, and file sizes into human-readable text formats, with support for fuzzy durations like "3 minutes ago" and localization to multiple languages.
Install it if you need to display human-readable numbers, durations, or sizes to end users.
Extract detailed information about text characters, rectangles, lines, and tables from machine-generated PDFs, with visual debugging support.
Not recommended for scanned image-based PDFs.
Parses and writes JSON5 data format, which extends JSON with JavaScript-style comments, unquoted keys, trailing commas, and single-quoted strings.
PrettyTable formats and displays tabular data as visually structured ASCII tables, with support for multiple output formats including HTML, JSON, CSV, LaTeX, and MediaWiki markup.
Install it if you need to display structured data in terminals, logs, or text-based reports.
Splits text documents into chunks using a variety of strategies, designed to work with LangChain's language model pipelines.
Parses dates from text in multiple languages and formats, handling both absolute dates and relative expressions like "3 days ago" or "next month".
Install it if you need to parse dates from unstructured text or HTML; skip it if you only work with standardized, single-format date strings.
Pyphen hyphenates text in multiple languages using Hunspell dictionaries included with the package, enabling proper word breaks for text layout and formatting.
Install it if you need reliable hyphenation for text layout or formatting.
Parses human-readable date and time strings (like "tomorrow" or "next Friday") into structured time values and Python datetime objects.
Parses human-readable time expressions (like '2h32m', '1.2 minutes', or '2 days, 4:13:02') into seconds as an integer or float.
Converts Unicode text to ASCII-safe representations by transliterating non-ASCII characters to their closest ASCII equivalents, useful for generating URL slugs, legacy system integration, and keyboard-friendly identifiers.
Install it if you need lossy Unicode-to-ASCII conversion and can accept GPLv2+ copyleft terms.
A Sphinx extension that outputs documentation in serialized HTML formats (JSON and pickle) alongside standard HTML.
A Sphinx extension that generates HTML help files from reStructuredText documentation source.
A Sphinx extension that converts documentation source files into QtHelp format, enabling integration with Qt-based help systems and documentation viewers.
However, be aware that maintenance is dormant—if you rely on it long-term, monitor Sphinx compatibility and consider whether QtHelp remains the right format for your…
Sphinx extension that generates Apple help books from reStructuredText documentation.
However, verify that Apple help format output remains compatible with your target macOS versions, and be aware that dormant maintenance means bug fixes or format…
A Sphinx extension that generates Devhelp documentation output format, allowing Sphinx-built documentation to be read in the Devhelp help viewer.
A Sphinx extension that renders mathematical display equations in HTML output using JavaScript.
natsort provides natural sorting for strings containing numbers, ordering them by numeric value rather than lexicographic order—so '10' comes after '9' instead of after '1'.
Install it if you need to sort strings containing numbers in a human-readable way.
Creates formatted ASCII tables for console output with support for alignment, data types, cell wrapping, and optional CJK text and emoji rendering via optional dependencies.
However, maintenance is dormant—no updates in over 1000 days—so consider it for stable, low-change use cases only.
Pytesseract wraps Google's Tesseract-OCR engine to extract text from images in Python, supporting multiple languages and output formats including plain text, bounding boxes, PDFs, and HOCR.
Levenshtein computes string edit distances, similarities, and related operations like approximate string averaging through a C extension for fast performance.
Install it if you need edit-distance operations and can accept the GPL-2.0-or-later copyleft requirement; avoid it if your project must remain proprietary without GPL…
Sanitize and validate filenames and file paths by removing or replacing invalid characters, reserved names, and unprintable characters across multiple platforms (Linux, Windows, macOS, POSIX, and universal).
Install it if you need reliable filename or path validation across platforms; the low friction and clear API make it a straightforward addition to any project…
Parses and serializes Amazon Ion, a self-describing data format that extends JSON with richer type support and binary encoding, using a Python API modeled on the standard library's json module.
MkDocs generates static HTML documentation sites from Markdown source files, configured via a single YAML file, with support for themes, plugins, and Markdown extensions.
Install it if you need to generate static documentation from Markdown with minimal configuration; skip it if you require cutting-edge feature development or need…
bashlex parses bash shell syntax into an abstract syntax tree (AST), supporting complex constructs like command and process substitutions that standard tools like shlex cannot handle.
A drop-in replacement for Python's re module that uses Google's RE2 engine for pattern matching, trading some PCRE features for guaranteed linear-time performance.
FuzzyWuzzy performs fuzzy string matching using Levenshtein distance to find approximate matches between text sequences, with optional speedup via an external C library.
However, do not adopt it for new projects requiring active maintenance or Python version support beyond 3.6.
Detects and fixes mojibake (garbled Unicode text caused by encoding mismatches) and recovers correctly-encoded text from multiple layers of encoding corruption.