Packages
-
charset-normalizer
Detects and normalizes text encoding from…
permissive · active · 1.7B/mo
-
tiktoken
tiktoken is a fast BPE tokenizer that converts…
permissive · active · 233.0M/mo
-
chardet
Detects character encoding and language in byte…
permissive · active · 199.0M/mo
-
text-unidecode
Converts Unicode text to ASCII by…
copyleft · abandoned · 89.0M/mo
-
lark
Lark is a parsing library that builds abstract…
permissive · active · 79.7M/mo
-
tree-sitter
Python bindings to the tree-sitter parsing…
permissive · active · 79.0M/mo
-
nltk
NLTK is a Python library for natural language…
permissive · active · 74.1M/mo
-
humanfriendly
Formats and parses numbers, file sizes,…
permissive · abandoned · 54.3M/mo
-
inflection
Inflection singularizes and pluralizes English…
permissive · abandoned · 51.7M/mo
-
snowballstemmer
Provides stemming algorithms for 34 languages,…
permissive · active · 36.1M/mo
-
sentencepiece
SentencePiece is an unsupervised text tokenizer…
permissive · active · 35.6M/mo
-
pyphen
Pyphen hyphenates text in multiple languages…
copyleft · active · 35.5M/mo
-
inflect
Generates correct English plurals, singulars,…
permissive · active · 31.6M/mo
-
tree-sitter-languages
Provides pre-compiled Python bindings for…
permissive · dormant · 31.2M/mo
-
tree-sitter-bash
Provides a tree-sitter grammar for parsing Bash…
permissive · aging · 26.1M/mo