Packages
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Provides stemming algorithms for 34 languages, reducing word variants to a common stem for text search and indexing applications.
Install it if you need stemming for search or text indexing.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
Converts HTML and CSS to PDF documents using Python, enabling developers with web skills to generate PDF templates without learning specialized PDF libraries.
However, the aging maintenance status (last release 537 days ago) warrants caution in production: verify that its rendering quality and feature set meet your specific…
Maps ISO 639 language codes (639-1, 639-2, 639-3) to language names and metadata, with methods to look up or guess a language from any code or name format.
StringZilla provides SIMD and SWAR-accelerated string operations including substring search, hashing, edit distances, sorting, and segmentation for Python, with no runtime dependencies.
Install only if your Python version is 3.10 or later.
Provides an async Python API to index and search HTML content, enabling full-text search indexing of static sites and custom content.
Install it if you need to index HTML or custom content for search—particularly valuable for static site generators and documentation platforms where you want search…
PyStemmer provides word stemming algorithms for multiple languages, reducing words to their morphological base form to improve search and information retrieval.
Jieba segments Chinese text into words using multiple algorithms (precise, full, and search-engine modes) and supports both simplified and traditional Chinese with custom dictionary injection.
However, install it only if you are working with Chinese text and can verify compatibility with your Python version—the package is dormant and may not work on very…
Extracts and cleans the main article text and title from HTML documents using lxml-based parsing and heuristic content detection.
A Python SDK for web scraping, crawling, searching, and extracting structured data from websites and research papers via the Firecrawl API, returning results as clean Markdown, HTML, or typed objects.
Whoosh is a pure-Python full-text search and indexing library that lets you add search functionality to applications without compiling native code.
OCRmyPDF adds searchable text layers to scanned PDF files using Tesseract OCR, enabling them to be searched and copy-pasted while optionally deskewing, cleaning, and converting to PDF/A format.
Install it if you need to make scanned PDFs searchable.
Computes Jaro and Jaro-Winkler string similarity scores, returning values from 0 (no match) to 1 (perfect match) for comparing two strings.
Provides a unified Python interface to download, parse, and iterate over information retrieval benchmarks and training datasets (documents, queries, relevance judgments) from many public sources.
Install it if you work with information retrieval datasets or need to prototype ranking models against standard benchmarks.
A Sphinx theme that extends PyData Sphinx Theme with NVIDIA branding and styling for documentation projects.
However, the proprietary license is restrictive—verify that your project qualifies as a 'NVIDIA product or service' before committing.
Computes phonetic keys of strings using Soundex, NYSIIS, Metaphone, and Double Metaphone algorithms for fuzzy matching and phonetic indexing.
Parses and crawls sitemaps in multiple formats (XML, RSS, Atom, plain text, Google News/Image) and extracts URLs efficiently without loading entire trees into memory.
Extends Python's set class to perform fuzzy string matching using N-gram similarity, allowing efficient searches for similar items in a collection.
However, do not adopt it for security-sensitive applications or if you require ongoing maintenance and updates.
PyTextRank implements graph-based TextRank and related algorithms as a spaCy pipeline extension to extract key phrases and perform extractive summarization on text documents.
CocoIndex maintains a live, incrementally-updated index of codebases, documents, and other sources for AI agents and LLM applications, recomputing only the changed portions rather than re-processing everything.
Whoosh-Reloaded is a pure-Python full-text indexing and search library that lets you add search functionality to applications without compiling native code.