Packages
Provides a Fortran language grammar for tree-sitter, enabling incremental parsing and syntax analysis of Fortran code.
PyThaiNLP provides Thai-language natural language processing tools including tokenization, part-of-speech tagging, transliteration, spelling correction, and linguistic utilities, designed as a Thai counterpart to NLTK.
Install it if you need to process Thai text; the base package is lightweight and the optional extras allow you to add machine translation or WordNet support as needed.
Segments provides Unicode-aware tokenization and orthography-based text segmentation, using orthography profiles to define how text should be split into meaningful units.
Install it if you need Unicode-aware tokenization with orthography profile support for linguistic or text-processing work; skip it if you only need basic whitespace…
Chonkie splits text into semantically meaningful chunks for RAG pipelines, offering multiple chunking strategies (recursive, semantic, token-based, code-aware) plus refinement and embedding integration.
VADER performs lexicon and rule-based sentiment analysis optimized for social media text, returning polarity scores for any input string.
Install it if you need quick, interpretable sentiment scores for social media or informal text without the overhead of training a model.
Counts syllables in English words, returning an integer for a given word string.
Install it if you need syllable counts in English text; the low friction and clear API make it practical for linguistic or readability analysis.
Provides string manipulation functions that work with grapheme clusters—user-perceived characters—rather than individual Unicode code points, enabling correct length calculations and slicing for strings with combining marks, emoji modifiers, and other multi-codepoint characters.
Spark NLP provides distributed natural language processing on Apache Spark, offering pretrained pipelines and models for tokenization, named entity recognition, sentiment analysis, machine translation, and embeddings across multiple languages.
Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.
OCRmyPDF adds searchable text layers to scanned PDF files using Tesseract OCR, enabling them to be searched and copy-pasted while optionally deskewing, cleaning, and converting to PDF/A format.
Install it if you need to make scanned PDFs searchable.
Splits concatenated words into their constituent parts using probabilistic English language modeling—for example, turning 'imateapot' into ['im', 'a', 'teapot'].
Converts English text to phoneme sequences using dictionary lookup, part-of-speech disambiguation, and neural prediction for out-of-vocabulary words.
Stanza is a Python NLP library that runs accurate natural language processing tools on 60+ languages, including tokenization, part-of-speech tagging, dependency parsing, and named entity recognition, with optional access to Java Stanford CoreNLP.
Install it if you need dependency parsing, NER, or POS tagging across many languages or in biomedical domains; skip it only if you need real-time performance on…
Python wrapper for MeCab, a morphological analyzer that tokenizes and tags Japanese text with part-of-speech information and lemmatization.
Loads the espeak-ng shared library and makes it available to other Python packages that need text-to-speech synthesis via espeak-ng.
However, verify the license terms first (currently unclear) and note that the package is dormant—no active maintenance since January 2025.
A Cython wrapper for MeCab that tokenizes and performs morphological analysis on Japanese text, returning parsed words with grammatical features.
Guesses the gender of a person from their first name using a lookup database, with support for country-specific inference and internationalized names.
Converts text to phonetic representations (phones) in multiple languages using pluggable backends like espeak, espeak-mbrola, festival, or segments.
Cython bindings for HarfBuzz that shape text using font files, converting character strings into positioned glyphs with kerning and ligature support.
pyttsx3 converts text to speech offline using your system's native TTS engines, with control over voice, speed, volume, and the ability to save audio to files.
Computes ROUGE scores (Recall-Oriented Understudy for Gisting Evaluation) to measure text summarization quality by comparing generated text against reference text.
However, verify that slight differences from official ROUGE-155 are acceptable for your use case, and be aware that no updates are forthcoming—use it as a stable…
Blingfire provides fast tokenization and text processing using finite state machines, supporting multiple algorithms (pattern-based, WordPiece, SentencePiece Unigram LM, BPE) with prebuilt models for BERT, XLNET, GPT-2, and other NLP frameworks.
However, be aware it is dormant (last release September 2021) and may not receive updates for new Python versions or model formats.
Nagisa performs Japanese word segmentation and part-of-speech tagging using recurrent neural networks, outputting tokenized words with their grammatical tags.
Parse, write, and transform Fluent localization files with built-in support for syntax trees, serialization, and tree traversal utilities.
A Python SDK for interacting with Smartling's translation API, enabling programmatic access to computer-assisted translation services and localization workflows.
However, verify API compatibility before use, and be aware that the aging maintenance status means you may need to maintain patches yourself if Smartling's API evolves.
Detects the language of text using Google's CLD2 library, supporting over 165 languages and returning reliability scores and byte-range vectors.
However, consider the aging maintenance status (last release 504 days ago): verify that it works reliably with your target Python version and that CLD2's detection…
Splits Indo-European text into sentences and words using rule-based segmentation and tokenization, with command-line tools for batch processing.
Unsupervised Korean natural language processing toolkit that extracts words and nouns from text, tokenizes sentences, and performs part-of-speech analysis without requiring training data.
However, GPLv3 licensing restricts proprietary use, and the package requires a reasonably large, homogeneous corpus to work well—single documents or mixed-domain text…
Provides Python bindings to Azure's cloud-based Natural Language Processing service for sentiment analysis, entity recognition, language detection, key phrase extraction, PII detection, and text summarization.
Install it if you need to call Azure's NLP service from Python; it is the canonical way to do so.
Provides the full-edition Sudachi dictionary resource for Japanese morphological analysis with SudachiPy, downloaded and installed as a Python package.
Identifies the language of text input across 97 languages using a pre-trained statistical model, available as a command-line tool, Python library, or WSGI web service.
However, do not use it if you require active maintenance, support for modern Python versions beyond basic compatibility, or confidence that the model reflects current…
Converts written text to phonetic representations (grapheme-to-phoneme conversion) for text-to-speech systems, supporting English, Japanese, Korean, Chinese, and Vietnamese with language-specific tokenization and phoneme rules.
However, the aging maintenance status (last release 496 days ago) means you should verify that the language and features you need are stable and that any issues may…
Converts natural language number words into their digit representations across seven languages, and detects ordinal, cardinal, and decimal numbers in text streams.
Converts numbers written in natural language (e.g., "twenty three") to their numeric equivalents, supporting cardinal numbers in English, Hindi, Spanish, Ukrainian, and Russian, plus English ordinals and fractions.
Converts text in any script to Latin alphabet romanization, supporting multiple languages with context-aware character mappings and optional language codes.
However, the dormant maintenance status (no activity for 777 days) means you should verify that its romanization rules and Unicode support remain adequate for your…
Jamo decomposes and synthesizes Hangul syllables into their constituent jamo (Korean letter components), enabling phonetic and spelling analysis of Korean text.
However, it is abandoned and untested on modern Python—verify it works in your environment before relying on it in production.
Generates random English words and sentences with filtering options (by length, starting/ending characters, part of speech, or regex), plus a command-line interface for quick generation.
Install it if you need random word or sentence generation; the low friction and permissive license make it a safe choice.
Kokoro is an inference library for the Kokoro-82M text-to-speech model, enabling you to generate spoken audio from text in multiple languages with a lightweight, open-weight neural network.
However, verify that the preset voices and supported languages meet your needs, and be aware that torch and transformers are heavy dependencies.
Zhon provides character constants and regular expression patterns for processing Chinese text, including CJK characters, Pinyin syllables, and Zhuyin marks.
Install it if you're building tools that need to identify, validate, or extract Chinese characters, Pinyin, or Zhuyin.
Detects and matches words that appear identical but use different Unicode characters, useful for finding homograph attacks, normalizing text, or bypassing character-based filters.
However, maintenance is dormant (last release 2021-03-24), so expect no updates; verify that Unicode version 8.0.0 is current enough for your threat model or data…
fastokens is a high-performance BPE tokenizer for large language models, built on a Rust backend and compatible with HuggingFace tokenizer.json and tiktoken model formats.
However, verify the license status in the repository before use in proprietary projects, and confirm that unsupported tokenizer features do not block your use case.