skillfed

Linguistic packages

262 packages · page 2 of 3

Packages

  • kokoro-onnx

    Converts text to speech using ONNX Runtime,…

    unclear · active · 550.9K/mo

  • hanzidentifier

    Identifies whether a string contains Simplified…

    permissive · dormant · 549.2K/mo

  • transliterate

    Converts text between Latin and non-Latin…

    copyleft · active · 545.7K/mo

  • llm

    LLM is a CLI tool and Python library for…

    permissive · active · 538.8K/mo

  • symspellpy

    symspellpy is a Python port of SymSpell v6.7.2…

    permissive · active · 518.4K/mo

  • tinysegmenter

    TinySegmenter is a compact Japanese tokenizer…

    permissive · abandoned · 510.1K/mo

  • unidic-lite

    Provides a pip-installable Japanese…

    permissive · abandoned · 494.3K/mo

  • tokie

    A fast, Rust-backed tokenizer library that…

    permissive · active · 487.3K/mo

  • OpenCC

    Converts text between Traditional Chinese,…

    permissive · active · 472.7K/mo

  • jieba3k

    Performs Chinese word segmentation, breaking…

    unclear · abandoned · 465.5K/mo

  • yake

    YAKE extracts keywords from text documents…

    copyleft · aging · 453.5K/mo

  • jinja2_pluralize

    Provides Jinja2 template filters for…

    permissive · abandoned · 443.4K/mo

  • py3langid

    Identifies the language of text in one of 97…

    permissive · dormant · 437.9K/mo

  • pymorphy3

    Morphological analyzer and POS tagger for…

    permissive · aging · 434.4K/mo

  • phonemizer

    Phonemizer converts written text into phonetic…

    copyleft · active · 431.8K/mo

  • pymorphy3-dicts-ru

    Provides Russian morphological dictionaries for…

    permissive · abandoned · 431.3K/mo

  • iso639-lang

    Resolves ISO 639 language codes and names to…

    permissive · aging · 423.3K/mo

  • unidic

    Provides the UniDic 2.3.0 Japanese…

    permissive · aging · 416.4K/mo

  • conllu

    Parses CoNLL-U formatted text (a standard NLP…

    unclear · active · 405.9K/mo

  • DAWG2-Python

    Read-only access to DAWG (directed acyclic word…

    permissive · active · 400.6K/mo

  • editdistpy

    Computes Levenshtein and Damerau-Levenshtein…

    permissive · active · 382.2K/mo

  • Janome

    Janome is a Japanese morphological analyzer…

    permissive · active · 380.3K/mo

  • azure-ai-translation-document

    Translates documents stored in Azure Blob…

    permissive · active · 377.3K/mo

  • sentence-stream

    Splits text streams into sentences even when…

    permissive · active · 375.2K/mo

  • mojimoji

    Converts Japanese text between hankaku…

    permissive · dormant · 375.1K/mo

  • cmudict

    Provides Python access to the CMU Pronouncing…

    copyleft · active · 359.3K/mo

  • language_tool_python

    Python wrapper for LanguageTool that checks…

    copyleft · active · 347.6K/mo

  • kiwipiepy

    Kiwipiepy tokenizes and analyzes Korean text…

    copyleft · active · 346.1K/mo

  • demoji

    demoji finds and removes emojis from text,…

    permissive · active · 338.2K/mo

  • azure-ai-translation-text

    Provides a Python client for Azure's neural…

    unclear · active · 335.7K/mo

  • rjieba

    A Python binding to the Rust-based jieba-rs…

    permissive · active · 334.7K/mo

  • phonetics

    Computes phonetic keys of strings using…

    permissive · abandoned · 328.6K/mo

  • alt-profanity-check

    Detects offensive or profane language in text…

    permissive · active · 324.9K/mo

  • mypy-boto3-comprehend

    Provides type annotations and IDE…

    permissive · active · 319.3K/mo

  • lorem

    Generates random placeholder text that…

    permissive · abandoned · 318.6K/mo

  • kaldialign

    Computes edit distance, alignment, and word…

    permissive · active · 318.6K/mo

  • clean-text

    Preprocesses and normalizes user-generated text…

    permissive · aging · 317.6K/mo

  • semantic-text-splitter

    Splits long text into semantically meaningful…

    permissive · active · 307.4K/mo

  • cn2an

    Converts between Chinese numerals and Arabic…

    permissive · active · 304.2K/mo

  • argostranslate

    Argos Translate is an open-source offline…

    permissive · active · 300.4K/mo

  • lemminflect

    Lemmatizes and inflects English words using…

    permissive · dormant · 296.8K/mo

  • normality

    Normality removes diacritics, punctuation, and…

    permissive · active · 294.8K/mo

  • wetext

    Normalizes and denormalizes text in Chinese,…

    permissive · active · 281.4K/mo

  • torchtext

    torchtext provides text datasets, preprocessing…

    permissive · abandoned · 279.9K/mo

  • nicknames

    Provides a curated dataset of English given…

    permissive · active · 276.9K/mo

  • edlib

    Edlib calculates edit distance (Levenshtein…

    permissive · aging · 273.6K/mo

  • Metaphone

    Implements the Metaphone and Double Metaphone…

    permissive · dormant · 272.7K/mo

  • emot

    Extracts emojis and emoticons from text,…

    unclear · dormant · 272.2K/mo

  • english

    English is a utility library providing English…

    permissive · abandoned · 251.4K/mo

  • udtools

    Validates CoNLL-U format files against…

    copyleft · active · 251.0K/mo

  • mwxml

    Efficiently parse and stream-process MediaWiki…

    permissive · active · 249.6K/mo

  • udapi

    Udapi is a Python framework for reading,…

    copyleft · active · 248.2K/mo

  • minisbd

    Detects sentence boundaries in text across many…

    agpl · active · 245.5K/mo

  • zhconv

    Converts text between Simplified and…

    copyleft · dormant · 243.4K/mo

  • text2digits

    Converts written-out number words in text to…

    permissive · active · 242.4K/mo

  • vale

    Vale is a command-line tool for enforcing…

    permissive · aging · 240.8K/mo

  • pynini

    Pynini compiles grammar rules into weighted…

    permissive · aging · 239.1K/mo

  • quantulum3

    Extracts quantities, measurements, and their…

    permissive · active · 226.9K/mo

  • simplemma

    Simplemma converts inflected word forms to…

    permissive · active · 220.5K/mo

  • DAWG-Python

    Reads DAWG (Directed Acyclic Word Graph) files…

    permissive · dormant · 219.9K/mo

  • unicode-rbnf

    Converts numbers to spelled-out text in…

    permissive · active · 216.5K/mo

  • pymorphy2

    Morphological analyzer and inflection engine…

    permissive · dormant · 215.9K/mo

  • ipadic

    Provides the IPAdic Japanese morphological…

    unclear · abandoned · 209.1K/mo

  • kiwipiepy-model

    Provides pre-trained morphological analysis…

    copyleft · active · 200.4K/mo

  • pymorphy2-dicts-ru

    Provides Russian morphological dictionaries for…

    permissive · dormant · 198.1K/mo

  • fasttext-langdetect

    Identifies the language of UTF-8 text using…

    permissive · active · 194.3K/mo

  • home-assistant-intents

    Provides intent recognition data and utilities…

    permissive · active · 192.5K/mo

  • hassil

    Parses natural language sentences into…

    permissive · active · 190.8K/mo

  • ngram

    Extends Python's set class to perform fuzzy…

    copyleft · abandoned · 180.8K/mo

  • sumy

    Sumy extracts summaries from HTML pages or…

    permissive · active · 173.5K/mo

  • pyvi

    Provides Vietnamese language processing tools…

    permissive · dormant · 166.0K/mo

  • panphon

    PanPhon maps International Phonetic Alphabet…

    permissive · aging · 161.7K/mo

  • wn

    Wn is a Python library for querying and…

    permissive · active · 157.4K/mo

  • konoha

    Konoha provides a unified Python interface to…

    permissive · active · 156.9K/mo

  • gladiaio-sdk

    A Python SDK for the Gladia speech-to-text API,…

    unclear · active · 154.9K/mo

  • pyxDamerauLevenshtein

    Computes Damerau-Levenshtein edit distance…

    permissive · active · 153.5K/mo

  • PyArabic

    PyArabic provides functions to manipulate…

    copyleft · abandoned · 152.3K/mo

  • translators

    Translators provides a unified Python interface…

    copyleft · aging · 147.2K/mo

  • hangul-romanize

    Converts Korean Hangul text to romanized (Latin…

    unclear · abandoned · 145.1K/mo

  • ctparse

    Parses natural language time expressions (like…

    permissive · abandoned · 143.9K/mo

  • rouge-metric

    Computes ROUGE metrics (ROUGE-N, ROUGE-L,…

    permissive · abandoned · 137.3K/mo

  • english-words

    Provides curated sets of English words from…

    permissive · aging · 136.8K/mo

  • mecab-ko

    Python wrapper for MeCab-ko, a morphological…

    permissive · aging · 135.5K/mo

  • konlpy

    KoNLPy provides Korean natural language…

    copyleft · abandoned · 134.8K/mo

  • wyoming

    Wyoming is a peer-to-peer TCP protocol for…

    permissive · active · 134.6K/mo

  • tree-sitter-hcl

    Provides a tree-sitter parser grammar for HCL…

    permissive · aging · 132.1K/mo

  • razdel

    Splits Russian text into sentences and tokens…

    permissive · active · 130.5K/mo

  • habachen

    Habachen converts between full-width and…

    permissive · aging · 126.8K/mo

  • textcase

    Converts strings between different text case…

    permissive · active · 125.7K/mo

  • cyrtranslit

    Converts text between Cyrillic and Latin…

    permissive · active · 125.3K/mo

  • bpemb

    BPEmb provides pre-trained subword embeddings…

    permissive · dormant · 124.6K/mo

  • pyleri

    Pyleri is a left-right parser generator that…

    permissive · active · 124.6K/mo

  • wordsegment

    Splits unsegmented English text into individual…

    permissive · abandoned · 123.7K/mo

  • tree-sitter-haskell

    Provides a tree-sitter parser grammar for…

    permissive · aging · 121.4K/mo

  • tokenizer

    Tokenizes Icelandic text into words,…

    permissive · active · 121.1K/mo

  • spanishconjugator

    Conjugates Spanish verbs by tense, mood, and…

    permissive · aging · 120.7K/mo

  • urduhack

    Urduhack provides NLP preprocessing,…

    permissive · dormant · 118.2K/mo

  • pytextrank

    PyTextRank implements graph-based TextRank and…

    permissive · active · 117.0K/mo

  • pymorphy3-dicts-uk

    Provides Ukrainian morphological dictionaries…

    permissive · abandoned · 113.8K/mo

  • names-dataset

    Looks up demographic information about first…

    permissive · aging · 112.9K/mo

  • nlptutti

    Measures Korean speech-to-text accuracy using…

    permissive · active · 112.7K/mo

  • cutlet

    Cutlet converts Japanese text to romaji (Latin…

    permissive · active · 110.4K/mo

  • gcld3

    Identifies the language of input text using a…

    unclear · abandoned · 109.8K/mo

  • orthography2ipa

    Converts spelling to IPA phonetic transcription…

    permissive · active · 107.3K/mo

  • epitran

    Epitran converts written text in various…

    permissive · active · 107.2K/mo

  • pypinyin-dict

    Provides alternative pinyin (romanized Chinese)…

    permissive · dormant · 105.3K/mo

  • rhoknp

    rhoknp is a Python binding for Japanese…

    permissive · active · 103.2K/mo

  • textacy

    textacy extends spaCy's NLP capabilities with…

    permissive · dormant · 102.8K/mo

  • g2pkk

    g2pkk converts Korean text to phonetic…

    permissive · abandoned · 102.4K/mo

  • gruut

    Gruut tokenizes, cleans, and converts text to…

    permissive · abandoned · 101.9K/mo

  • SudachiDict-small

    Provides the small-edition Sudachi dictionary…

    permissive · active · 99.9K/mo

  • mecab-ko-dic

    Provides a Korean dictionary for MeCab…

    unclear · abandoned · 99.2K/mo

  • whisper-timestamped

    Adds word-level timestamps and confidence…

    copyleft · aging · 98.9K/mo

  • indic-nlp-library

    Indic NLP Library provides text processing and…

    permissive · dormant · 98.7K/mo

  • lexical-diversity

    Calculates lexical diversity metrics (TTR,…

    permissive · dormant · 98.7K/mo

  • chompjs

    Parses JavaScript objects embedded in HTML or…

    permissive · active · 98.1K/mo

  • gruut-ipa

    Gruut IPA parses, analyzes, and converts…

    permissive · abandoned · 98.1K/mo

  • gruut-lang-en

    Provides English language data files for…

    permissive · abandoned · 96.8K/mo

  • indic-transliteration

    Converts text between different Indic script…

    permissive · active · 95.2K/mo

  • deepsearch-glm

    Extracts entities, relations, and linguistic…

    permissive · dormant · 92.2K/mo