Packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Formats and parses numbers, file sizes, timespans, and other values into human-readable text; provides terminal interaction utilities including ANSI text styling and user prompts.
Inflection singularizes and pluralizes English words, and converts strings between CamelCase and underscored formats.
Provides stemming algorithms for 34 languages, reducing word variants to a common stem for text search and indexing applications.
Install it if you need stemming for search or text indexing.
SentencePiece is an unsupervised text tokenizer and detokenizer that converts text into token IDs or subword pieces, and reconstructs text from tokens. It supports single and batch encoding/decoding with multiple output formats including NumPy arrays, offset mappings, and Protobuf messages.
Pyphen hyphenates text in multiple languages using Hunspell dictionaries included with the package, enabling proper word breaks for text layout and formatting.
Install it if you need reliable hyphenation for text layout or formatting.
Generates correct English plurals, singulars, ordinals, indefinite articles, and word-based number representations through a rule-based engine.
Install it if your application generates user-facing English text that must agree in number or convert numbers to words.
Provides pre-compiled Python bindings for tree-sitter language parsers, eliminating the need to download and build language support individually.
However, the package is dormant (latest release 2024-02-04) and depends on tree-sitter; verify that the bundled language versions and tree-sitter compatibility match…
Provides a tree-sitter grammar for parsing Bash shell scripts into an abstract syntax tree, enabling programmatic analysis and manipulation of Bash code.
Provides a JavaScript and JSX parser grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of JavaScript code.
Extracts original and updated publication dates from web pages by parsing HTML markup, metadata, and text content, with both Python API and command-line interfaces.
Install it if you need reliable date extraction from web pages; the fast mode offers good speed and the extensive mode provides high recall when accuracy matters most.
Trafilatura extracts main text, metadata, and structured content from web pages and HTML, converting raw HTML into clean, usable data in multiple output formats.
Validates, normalizes, filters, and samples URLs for web crawling and document collection, removing spam, trackers, and low-value pages while respecting language and content-type constraints.
Install it if you are building a crawler or bulk collection pipeline and need to filter, normalize, or sample URLs at scale.
Converts text to speech using Microsoft Edge's online TTS service, available as both a Python module and command-line tools for generating audio files and subtitles.
Parses, validates, and standardizes IETF language tags (BCP 47 format), converting between different representations of the same language and resolving deprecated or non-standard codes.
Install it if you work with language codes, multilingual systems, or need to normalize IETF language tags.
Detects the language of text input, supporting 55 languages and returning either a language code or probability scores for candidate languages.
Provides a C language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of C code.
Provides a Java language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of Java source code.
Provides a C# language parser for tree-sitter, enabling incremental syntax analysis and parsing of C# code from versions 1 through 13.0.
Install it if you're building tools that parse or analyze C# code and want fast, dependency-light syntax tree generation.
Provides a Rust language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of Rust code.
Provides a Go language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of Go source code.
Provides a Python grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of Python code.
However, the last release was 337 days ago—verify that the grammar version meets your Python syntax requirements before committing to it in production.
Converts numbers to their word representations in multiple languages, supporting cardinal numbers, ordinals, years, and currency formats.
Install it if you need multilingual number-to-words conversion.
Gensim is a Python library for topic modeling, document indexing, and similarity retrieval on large text corpora, using algorithms like LDA, LSA, word2vec, and others.
However, the aging maintenance status (300 days since last release) suggests the project is in steady-state rather than actively developed; evaluate whether its…
Provides TypeScript and TSX language grammars for tree-sitter, enabling incremental parsing and syntax tree generation for TypeScript and JSX code.
TextBlob provides a simple API for common natural language processing tasks including sentiment analysis, part-of-speech tagging, noun phrase extraction, tokenization, and text classification.
Provides a PHP language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of PHP code.
Provides a Ruby language grammar for tree-sitter, enabling incremental parsing and abstract syntax tree generation for Ruby code.
polib reads, writes, and manipulates gettext translation files (PO, POT, and MO formats), allowing you to load, iterate, modify entries, and create translation catalogs programmatically.
Install it if you need to work with gettext files programmatically.
Python binding to CRFsuite for conditional random field sequence labeling and structured prediction.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
A tree-sitter parser grammar for YAML files that enables incremental parsing and syntax tree extraction from YAML 1.2 documents.
Not recommended if you simply need to read or write YAML data—use PyYAML or similar instead.
Detects sentence boundaries in text using rule-based heuristics, splitting paragraphs into individual sentences while handling edge cases like abbreviations and decimal points.
However, do not install if you need active maintenance, multi-language support, or assurance of ongoing bug fixes.
Provides fast, parallel word stemming using Snowball algorithms via a Rust backend, with methods for single words, sequential lists, and parallel batch processing.