Packages
-
kokoro-onnx
Converts text to speech using ONNX Runtime,…
unclear · active · 550.9K/mo
-
hanzidentifier
Identifies whether a string contains Simplified…
permissive · dormant · 549.2K/mo
-
transliterate
Converts text between Latin and non-Latin…
copyleft · active · 545.7K/mo
-
llm
LLM is a CLI tool and Python library for…
permissive · active · 538.8K/mo
-
symspellpy
symspellpy is a Python port of SymSpell v6.7.2…
permissive · active · 518.4K/mo
-
tinysegmenter
TinySegmenter is a compact Japanese tokenizer…
permissive · abandoned · 510.1K/mo
-
unidic-lite
Provides a pip-installable Japanese…
permissive · abandoned · 494.3K/mo
-
tokie
A fast, Rust-backed tokenizer library that…
permissive · active · 487.3K/mo
-
OpenCC
Converts text between Traditional Chinese,…
permissive · active · 472.7K/mo
-
jieba3k
Performs Chinese word segmentation, breaking…
unclear · abandoned · 465.5K/mo
-
yake
YAKE extracts keywords from text documents…
copyleft · aging · 453.5K/mo
-
jinja2_pluralize
Provides Jinja2 template filters for…
permissive · abandoned · 443.4K/mo
-
py3langid
Identifies the language of text in one of 97…
permissive · dormant · 437.9K/mo
-
pymorphy3
Morphological analyzer and POS tagger for…
permissive · aging · 434.4K/mo
-
phonemizer
Phonemizer converts written text into phonetic…
copyleft · active · 431.8K/mo
-
pymorphy3-dicts-ru
Provides Russian morphological dictionaries for…
permissive · abandoned · 431.3K/mo
-
iso639-lang
Resolves ISO 639 language codes and names to…
permissive · aging · 423.3K/mo
-
unidic
Provides the UniDic 2.3.0 Japanese…
permissive · aging · 416.4K/mo
-
conllu
Parses CoNLL-U formatted text (a standard NLP…
unclear · active · 405.9K/mo
-
DAWG2-Python
Read-only access to DAWG (directed acyclic word…
permissive · active · 400.6K/mo
-
editdistpy
Computes Levenshtein and Damerau-Levenshtein…
permissive · active · 382.2K/mo
-
Janome
Janome is a Japanese morphological analyzer…
permissive · active · 380.3K/mo
-
azure-ai-translation-document
Translates documents stored in Azure Blob…
permissive · active · 377.3K/mo
-
sentence-stream
Splits text streams into sentences even when…
permissive · active · 375.2K/mo
-
mojimoji
Converts Japanese text between hankaku…
permissive · dormant · 375.1K/mo
-
cmudict
Provides Python access to the CMU Pronouncing…
copyleft · active · 359.3K/mo
-
language_tool_python
Python wrapper for LanguageTool that checks…
copyleft · active · 347.6K/mo
-
kiwipiepy
Kiwipiepy tokenizes and analyzes Korean text…
copyleft · active · 346.1K/mo
-
demoji
demoji finds and removes emojis from text,…
permissive · active · 338.2K/mo
-
azure-ai-translation-text
Provides a Python client for Azure's neural…
unclear · active · 335.7K/mo
-
rjieba
A Python binding to the Rust-based jieba-rs…
permissive · active · 334.7K/mo
-
phonetics
Computes phonetic keys of strings using…
permissive · abandoned · 328.6K/mo
-
alt-profanity-check
Detects offensive or profane language in text…
permissive · active · 324.9K/mo
-
mypy-boto3-comprehend
Provides type annotations and IDE…
permissive · active · 319.3K/mo
-
lorem
Generates random placeholder text that…
permissive · abandoned · 318.6K/mo
-
kaldialign
Computes edit distance, alignment, and word…
permissive · active · 318.6K/mo
-
clean-text
Preprocesses and normalizes user-generated text…
permissive · aging · 317.6K/mo
-
semantic-text-splitter
Splits long text into semantically meaningful…
permissive · active · 307.4K/mo
-
cn2an
Converts between Chinese numerals and Arabic…
permissive · active · 304.2K/mo
-
argostranslate
Argos Translate is an open-source offline…
permissive · active · 300.4K/mo
-
lemminflect
Lemmatizes and inflects English words using…
permissive · dormant · 296.8K/mo
-
normality
Normality removes diacritics, punctuation, and…
permissive · active · 294.8K/mo
-
wetext
Normalizes and denormalizes text in Chinese,…
permissive · active · 281.4K/mo
-
torchtext
torchtext provides text datasets, preprocessing…
permissive · abandoned · 279.9K/mo
-
nicknames
Provides a curated dataset of English given…
permissive · active · 276.9K/mo
-
edlib
Edlib calculates edit distance (Levenshtein…
permissive · aging · 273.6K/mo
-
Metaphone
Implements the Metaphone and Double Metaphone…
permissive · dormant · 272.7K/mo
-
emot
Extracts emojis and emoticons from text,…
unclear · dormant · 272.2K/mo
-
english
English is a utility library providing English…
permissive · abandoned · 251.4K/mo
-
udtools
Validates CoNLL-U format files against…
copyleft · active · 251.0K/mo
-
mwxml
Efficiently parse and stream-process MediaWiki…
permissive · active · 249.6K/mo
-
udapi
Udapi is a Python framework for reading,…
copyleft · active · 248.2K/mo
-
minisbd
Detects sentence boundaries in text across many…
agpl · active · 245.5K/mo
-
zhconv
Converts text between Simplified and…
copyleft · dormant · 243.4K/mo
-
text2digits
Converts written-out number words in text to…
permissive · active · 242.4K/mo
-
vale
Vale is a command-line tool for enforcing…
permissive · aging · 240.8K/mo
-
pynini
Pynini compiles grammar rules into weighted…
permissive · aging · 239.1K/mo
-
quantulum3
Extracts quantities, measurements, and their…
permissive · active · 226.9K/mo
-
simplemma
Simplemma converts inflected word forms to…
permissive · active · 220.5K/mo
-
DAWG-Python
Reads DAWG (Directed Acyclic Word Graph) files…
permissive · dormant · 219.9K/mo
-
unicode-rbnf
Converts numbers to spelled-out text in…
permissive · active · 216.5K/mo
-
pymorphy2
Morphological analyzer and inflection engine…
permissive · dormant · 215.9K/mo
-
ipadic
Provides the IPAdic Japanese morphological…
unclear · abandoned · 209.1K/mo
-
kiwipiepy-model
Provides pre-trained morphological analysis…
copyleft · active · 200.4K/mo
-
pymorphy2-dicts-ru
Provides Russian morphological dictionaries for…
permissive · dormant · 198.1K/mo
-
fasttext-langdetect
Identifies the language of UTF-8 text using…
permissive · active · 194.3K/mo
-
home-assistant-intents
Provides intent recognition data and utilities…
permissive · active · 192.5K/mo
-
hassil
Parses natural language sentences into…
permissive · active · 190.8K/mo
-
ngram
Extends Python's set class to perform fuzzy…
copyleft · abandoned · 180.8K/mo
-
sumy
Sumy extracts summaries from HTML pages or…
permissive · active · 173.5K/mo
-
pyvi
Provides Vietnamese language processing tools…
permissive · dormant · 166.0K/mo
-
panphon
PanPhon maps International Phonetic Alphabet…
permissive · aging · 161.7K/mo
-
wn
Wn is a Python library for querying and…
permissive · active · 157.4K/mo
-
konoha
Konoha provides a unified Python interface to…
permissive · active · 156.9K/mo
-
gladiaio-sdk
A Python SDK for the Gladia speech-to-text API,…
unclear · active · 154.9K/mo
-
pyxDamerauLevenshtein
Computes Damerau-Levenshtein edit distance…
permissive · active · 153.5K/mo
-
PyArabic
PyArabic provides functions to manipulate…
copyleft · abandoned · 152.3K/mo
-
translators
Translators provides a unified Python interface…
copyleft · aging · 147.2K/mo
-
hangul-romanize
Converts Korean Hangul text to romanized (Latin…
unclear · abandoned · 145.1K/mo
-
ctparse
Parses natural language time expressions (like…
permissive · abandoned · 143.9K/mo
-
rouge-metric
Computes ROUGE metrics (ROUGE-N, ROUGE-L,…
permissive · abandoned · 137.3K/mo
-
english-words
Provides curated sets of English words from…
permissive · aging · 136.8K/mo
-
mecab-ko
Python wrapper for MeCab-ko, a morphological…
permissive · aging · 135.5K/mo
-
konlpy
KoNLPy provides Korean natural language…
copyleft · abandoned · 134.8K/mo
-
wyoming
Wyoming is a peer-to-peer TCP protocol for…
permissive · active · 134.6K/mo
-
tree-sitter-hcl
Provides a tree-sitter parser grammar for HCL…
permissive · aging · 132.1K/mo
-
razdel
Splits Russian text into sentences and tokens…
permissive · active · 130.5K/mo
-
habachen
Habachen converts between full-width and…
permissive · aging · 126.8K/mo
-
textcase
Converts strings between different text case…
permissive · active · 125.7K/mo
-
cyrtranslit
Converts text between Cyrillic and Latin…
permissive · active · 125.3K/mo
-
bpemb
BPEmb provides pre-trained subword embeddings…
permissive · dormant · 124.6K/mo
-
pyleri
Pyleri is a left-right parser generator that…
permissive · active · 124.6K/mo
-
wordsegment
Splits unsegmented English text into individual…
permissive · abandoned · 123.7K/mo
-
tree-sitter-haskell
Provides a tree-sitter parser grammar for…
permissive · aging · 121.4K/mo
-
tokenizer
Tokenizes Icelandic text into words,…
permissive · active · 121.1K/mo
-
spanishconjugator
Conjugates Spanish verbs by tense, mood, and…
permissive · aging · 120.7K/mo
-
urduhack
Urduhack provides NLP preprocessing,…
permissive · dormant · 118.2K/mo
-
pytextrank
PyTextRank implements graph-based TextRank and…
permissive · active · 117.0K/mo
-
pymorphy3-dicts-uk
Provides Ukrainian morphological dictionaries…
permissive · abandoned · 113.8K/mo
-
names-dataset
Looks up demographic information about first…
permissive · aging · 112.9K/mo
-
nlptutti
Measures Korean speech-to-text accuracy using…
permissive · active · 112.7K/mo
-
cutlet
Cutlet converts Japanese text to romaji (Latin…
permissive · active · 110.4K/mo
-
gcld3
Identifies the language of input text using a…
unclear · abandoned · 109.8K/mo
-
orthography2ipa
Converts spelling to IPA phonetic transcription…
permissive · active · 107.3K/mo
-
epitran
Epitran converts written text in various…
permissive · active · 107.2K/mo
-
pypinyin-dict
Provides alternative pinyin (romanized Chinese)…
permissive · dormant · 105.3K/mo
-
rhoknp
rhoknp is a Python binding for Japanese…
permissive · active · 103.2K/mo
-
textacy
textacy extends spaCy's NLP capabilities with…
permissive · dormant · 102.8K/mo
-
g2pkk
g2pkk converts Korean text to phonetic…
permissive · abandoned · 102.4K/mo
-
gruut
Gruut tokenizes, cleans, and converts text to…
permissive · abandoned · 101.9K/mo
-
SudachiDict-small
Provides the small-edition Sudachi dictionary…
permissive · active · 99.9K/mo
-
mecab-ko-dic
Provides a Korean dictionary for MeCab…
unclear · abandoned · 99.2K/mo
-
whisper-timestamped
Adds word-level timestamps and confidence…
copyleft · aging · 98.9K/mo
-
indic-nlp-library
Indic NLP Library provides text processing and…
permissive · dormant · 98.7K/mo
-
lexical-diversity
Calculates lexical diversity metrics (TTR,…
permissive · dormant · 98.7K/mo
-
chompjs
Parses JavaScript objects embedded in HTML or…
permissive · active · 98.1K/mo
-
gruut-ipa
Gruut IPA parses, analyzes, and converts…
permissive · abandoned · 98.1K/mo
-
gruut-lang-en
Provides English language data files for…
permissive · abandoned · 96.8K/mo
-
indic-transliteration
Converts text between different Indic script…
permissive · active · 95.2K/mo
-
deepsearch-glm
Extracts entities, relations, and linguistic…
permissive · dormant · 92.2K/mo