$npx skillfedfor your agent

nlptutti

Korean STT CER, WER, and CRR metrics with explicit normalization and reproducible reports

Worth itPyPI Artificial IntelligenceReleased Aug 2026112.7K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — nlptutti-0.0.0.19-py3-none-any.whl
v0.0.0.19 · released 2026-08-12 · Python >=3.8 · 1 runtime deps: jiwer

Yes. Active, recently released package with low install friction, permissive MIT license, no known vulnerabilities, and a focused scope solving a real problem for Korean STT evaluation. The single dependency (jiwer) is lightweight. Best for teams evaluating Korean speech recognition systems or comparing STT services; less relevant if you work exclusively with non-Korean languages.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low friction: pure Python wheel with a single runtime dependency (jiwer).
  • Active maintenance with recent release (2 days old) and 73 repository stars.
  • Supports Python 3.8 through 3.14.

License · maintenance · safety

MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

last release 2026-08-12 (2 days) · last repo commit 2026-08-12 · 73 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 112,676 downloads/mo, #12,362 on PyPI

Verify before relying

pip install nlptutti

import nlptutti as metrics

reference = "오늘 날씨가 맑습니다"
hypothesis = "오늘 날씨는 맑습니다"

cer = metrics.get_cer(reference, hypothesis, rate_mode="standard")
print(round(cer["cer"], 4), cer["substitutions"])
  • Whether jiwer's Levenshtein distance implementation handles Korean combining characters correctly without explicit NFC normalization
  • Performance characteristics with large corpora (corpus size thresholds for practical use)
Same gist for agents: .md · .json

What it is and what it does

nlptutti is a Python package for evaluating Korean speech-to-text output against reference transcripts. It computes character error rate (CER), word error rate (WER), character correct rate (CRR), and entity/keyword preservation metrics by calculating substitutions, deletions, and insertions from Levenshtein minimum edit distance. The package supports both plain strings and structured formats (JSON, SRT, TSV) and offers two calculation modes: a normalized formula (default, for backward compatibility) and a standard formula (for model comparisons and published benchmarks).

The package handles Korean-specific concerns including punctuation removal, Unicode normalization options for combining characters, and word-boundary handling across different tokenization policies. It can evaluate single sentence pairs, entire corpora with micro/macro aggregation, named entities and keywords against provided dictionaries, and error patterns. Results include detailed edit operation counts and optional provenance tracking (SHA-256, package version, actual calculation parameters) for reproducibility.

Use it for

  • Compare accuracy of cloud STT services (Azure Speech, Amazon Transcribe, Google Cloud Speech-to-Text) using a consistent Korean-aware metric
  • Evaluate open-source speech recognition models (Whisper, FunASR) against reference transcripts with standard CER/WER formulas
  • Track preservation of named entities and product names in STT output using entity-specific CER and F1 scores
  • Reproduce historical evaluation results by maintaining the default normalized calculation mode across versions
  • Analyze error patterns (substitutions, deletions, insertions) to diagnose STT failure modes in Korean text

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Active, recently released package with low install friction, permissive MIT license, no known vulnerabilities, and a focused scope solving a real problem for Korean STT evaluation. The single dependency (jiwer) is lightweight. Best for teams evaluating Korean speech recognition systems or comparing STT services; less relevant if you work exclusively with non-Korean languages.

Install

nlptutti on PyPI

Before you install

Low friction: pure Python wheel with a single runtime dependency (jiwer). Active maintenance with recent release (2 days old) and 73 repository stars. Supports Python 3.8 through 3.14.

License in practice

MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

Quickstart

pip install nlptutti

import nlptutti as metrics

reference = "오늘 날씨가 맑습니다"
hypothesis = "오늘 날씨는 맑습니다"

cer = metrics.get_cer(reference, hypothesis, rate_mode="standard")
print(round(cer["cer"], 4), cer["substitutions"])

Verify before relying

  • Whether jiwer's Levenshtein distance implementation handles Korean combining characters correctly without explicit NFC normalization
  • Performance characteristics with large corpora (corpus size thresholds for practical use)

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
jiwer
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads112,676 / month, #12,362 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Text Processing :: Linguistic

Evidence: nlptutti-0.0.0.19-py3-none-any.whl

Tags

Capabilities
korean stt evaluation metricscer wer measurementspeech recognition error ratekorean transcription qualityasr performance metricskorean nlp evaluationtranscript accuracy testing
Topics
korean-nlpspeech-recognitionevaluation-metrics
PyPI keywords
STTASRKoreanNLPCERWERCRRKorean text normalizationUnicode normalizationspeech recognitiontranscript evaluationJSONSRTTSVFunASR

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “korean stt evaluation metrics”

  • nlptuttiMeasures Korean speech-to-text accuracy using character error rate…
  • pykrxScrapes Korean stock and bond market data from KRX and Naver,…
  • livekit-plugins-assemblyaiIntegrates AssemblyAI speech-to-text into the LiveKit Agents…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also texterrors · jiwer · kaldialign · openai-whisper · whisper-timestamped · SpeechRecognition · rouge · azure-cognitiveservices-speech · rouge-chinese · google-cloud-speech