Packages
Provides pre-compiled tree-sitter parsers for 371 programming languages, enabling fast syntax tree parsing and code analysis without requiring grammar compilation.
Provides a C++ language grammar for tree-sitter, enabling incremental parsing and syntax tree generation for C++ code.
Provides memory-efficient trie data structures for fast string lookups and prefix searches, using up to 50x-100x less memory than standard Python dicts while maintaining comparable lookup speed.
Detects the language of text using FastText models, returning language codes and confidence scores with minimal dependencies and offline capability.
Provides multilingual language names, speaker populations, and writing population data for standardized language codes, designed as a data supplement to the langcodes module.
Maps ISO 639 language codes (639-1, 639-2, 639-3) to language names and metadata, with methods to look up or guess a language from any code or name format.
Converts Unicode text to ASCII-only equivalents by transliterating each character to its best ASCII representation, leaving ASCII input unchanged and removing unmapped characters.
PyEnchant provides Python bindings to the Enchant spellchecking library, enabling spell-checking and correction suggestions across multiple languages and dictionary backends.
Lark is a parsing library that builds abstract syntax trees from context-free grammars using Earley, LALR(1), or CYK parsing algorithms, with minimal code required.
Provides a JSON grammar parser for tree-sitter, enabling incremental parsing and syntax tree analysis of JSON documents.
Provides a tree-sitter parser for embedded template languages like ERB and EJS, enabling syntactic analysis of code embedded within text using `<%` and `%>` delimiters.
Parses human names into seven structured fields—title, given, middle, family, suffix, nickname, maiden—with support for multiple scripts including East Asian names, and provides immutable results with composable configuration.
Install it if you need to parse or normalize human names, especially in multilingual contexts.
Sacremoses provides tokenization, detokenization, truecasing, and punctuation normalization for text processing, primarily for machine translation and NLP pipelines.
Install it if you need Moses-style tokenization or truecasing for machine translation or similar tasks.
Detects which language a text is written in, supporting 75 languages with high accuracy on both short snippets and full sentences using compiled Rust bindings.
PyStemmer provides word stemming algorithms for multiple languages, reducing words to their morphological base form to improve search and information retrieval.
Translates text between languages using multiple free translation services (Google, Microsoft, DeepL, Yandex, and others) with automatic language detection and batch processing support.
However, do not rely on it for production systems where translation service availability is critical—the dormant status means breakage from service changes will not…
Jieba segments Chinese text into words using multiple algorithms (precise, full, and search-engine modes) and supports both simplified and traditional Chinese with custom dictionary injection.
However, install it only if you are working with Chinese text and can verify compatibility with your Python version—the package is dormant and may not work on very…
wordfreq looks up word frequencies across over 40 languages from multiple corpus sources, returning frequencies as decimal values or on a logarithmic Zipf scale.
Provides a Kotlin language grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of Kotlin code.
However, maintenance is dormant (last release 582 days ago), so verify that the grammar covers your Kotlin version and feature set before committing to a production…
SudachiPy is a Python binding for Sudachi.rs, a Japanese morphological analyzer that tokenizes Japanese text into morphemes with part-of-speech tags, normalized forms, and reading information.
The main gotcha is the separate dictionary dependency and medium install friction from compiled wheels, but prebuilt binaries cover common platforms.
Provides an HTML grammar for tree-sitter, enabling incremental parsing and syntax tree analysis of HTML documents.
Provides a tree-sitter parser grammar for TOML files, enabling incremental parsing and syntax tree analysis of TOML configuration documents.
Install only if you already use or plan to use tree-sitter's Python API; this package alone does not parse TOML.
Provides a tree-sitter parser grammar for XML and DTD files, enabling incremental parsing and syntax tree analysis of XML documents.
Install only if you already use or plan to use tree-sitter's Python bindings.
Provides a tree-sitter SQL grammar for parsing SQL code into an abstract syntax tree, enabling programmatic analysis and manipulation of SQL statements.
Provides string manipulation functions that work with user-perceived characters (grapheme clusters) as defined by Unicode Standard Annex #29, rather than individual Unicode code points.
Provides the core edition of the Sudachi morphological analyzer dictionary for use with SudachiPy, enabling Japanese text tokenization and linguistic analysis.
Provides a tree-sitter grammar for parsing Markdown documents, enabling syntactic analysis of Markdown text according to CommonMark with optional extensions like GitHub flavored markdown, task lists, and tables.
However, do not use it where Markdown correctness is critical—the authors acknowledge known inaccuracies.
Provides a CSS grammar parser for tree-sitter, enabling incremental parsing and syntax tree analysis of CSS code.
Provides a tree-sitter grammar for parsing regular expressions in PCRE2, POSIX, and JavaScript syntaxes, enabling incremental parsing and syntax analysis of regex patterns.
However, it is a specialized grammar module—install it only if you have a specific need for tree-sitter-based regex parsing; it is not a general-purpose regex library.
Provides a tree-sitter grammar for parsing Swift source code, enabling incremental syntax analysis and abstract syntax tree generation.
Provides an Elixir language grammar for tree-sitter, enabling incremental parsing and syntax analysis of Elixir source code.
Provides a tree-sitter grammar for parsing Scala 2 and Scala 3 source code, enabling incremental syntax analysis and code structure extraction.
Install only if your target platform has a precompiled wheel available.
Provides a tree-sitter grammar for parsing Lua 5.x and LuaJIT 2.x code, enabling incremental syntax analysis and AST generation.
oslo.i18n provides utilities for handling internationalization and translation of text strings in Python applications, with support for multiple languages and locales.
Provides a Groovy language grammar for tree-sitter, enabling incremental parsing and syntax analysis of Groovy code.
However, dormant maintenance (no activity since 2024-11-19) and the author's own caveat that it is unpolished mean you should verify the grammar covers your use case…
Provides a tree-sitter grammar for parsing PowerShell code, enabling incremental syntax analysis and tree-based code inspection for PowerShell scripts.
Provides a tree-sitter grammar for parsing Objective-C code into syntax trees, enabling incremental parsing and code analysis.
Provides a tree-sitter grammar for parsing Zig source code, enabling incremental syntax analysis and AST generation for Zig programs.
Provides incremental parsing of Julia source code using tree-sitter, producing syntax trees suitable for code analysis and tooling.
Provides a SystemVerilog grammar parser for tree-sitter, enabling incremental parsing and syntax tree analysis of Verilog hardware description language code.