misaki
G2P engine for TTS
Decision gist · record as of 2026-08-14
Yes, if you need multilingual grapheme-to-phoneme conversion for text-to-speech. The package is permissively licensed under Apache License Version 2.0, has low install friction, and covers five languages with specialized tokenization. However, the aging maintenance status (last release 496 days ago) means you should verify that the language and features you need are stable and that any issues may see slow response.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Language-specific extras (e.g., [en], [ja], [ko], [zh], [vi]) must be installed; optional fallback to espeak requires system-level installation of espeak-ng.
- Low install friction with only two runtime dependencies (addict and regex).
- Maintenance status is aging—last release was 496 days ago—so expect slower response to issues, though the package appears stable for its current use case.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache License Version 2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions; suitable for both open-source and commercial projects provided you include license attribution.
last release 2025-04-05 (496 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 688,566 downloads/mo, #5,340 on PyPI
Alternatives
Verify before relying
pip install "misaki[en]"
from misaki import en
g2p = en.G2P(trf=False, british=False, fallback=None)
text = 'Text to convert to phonemes.'
phonemes, tokens = g2p(text)
print(phonemes)- Performance characteristics (latency, throughput) for typical text lengths and batch sizes.
- Accuracy metrics or benchmarks against reference datasets.
- Whether transformer-based mode (trf=True) requires additional dependencies or model downloads.
- Compatibility and behavior with out-of-vocabulary or mixed-language text.
- Current state of language-specific features (pitch accent for Japanese, etc.) relative to the description.
What it is and what it does
This package is a grapheme-to-phoneme (G2P) engine that takes written text and produces phonetic transcriptions in IPA notation for text-to-speech synthesis. It supports five languages (English, Japanese, Korean, Chinese, Vietnamese) with language-specific tokenization pipelines: English uses spaCy and num2words; Japanese leverages pyopenjtalk with pitch accent support; Korean adapts g2pkc; Chinese uses jieba and pinyin conversion; Vietnamese relies on Viphoneme.
The core workflow is simple: instantiate a language-specific G2P object, pass text through it, and receive both phoneme sequences and token alignments. You can run without transformers for speed or enable transformer-based processing for context-aware disambiguation. Optional fallback to espeak handles out-of-vocabulary words. With only two runtime dependencies (addict and regex) and a pure-Python wheel, installation is straightforward, though language-specific extras must be selected at install time.
Use it for
- Build multilingual TTS pipelines by converting text to phonemes for speech synthesis models.
- Preprocess text corpora for phonetic analysis or linguistic research across multiple languages.
- Handle out-of-vocabulary words in speech synthesis by falling back to espeak when dictionary lookup fails.
- Disambiguate homographs using optional transformer-based context when trf=True.
- Generate phonetic training data for speech models by batch-converting text documents to aligned phoneme sequences.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need multilingual grapheme-to-phoneme conversion for text-to-speech.
The package is permissively licensed under Apache License Version 2.0, has low install friction, and covers five languages with specialized tokenization. However, the aging maintenance status (last release 496 days ago) means you should verify that the language and features you need are stable and that any issues may see slow response.
Install
misaki on PyPI
Before you install
Low install friction with only two runtime dependencies (addict and regex). Maintenance status is aging—last release was 496 days ago—so expect slower response to issues, though the package appears stable for its current use case.
Language-specific extras (e.g., [en], [ja], [ko], [zh], [vi]) must be installed; optional fallback to espeak requires system-level installation of espeak-ng.
License in practice
Licensed under Apache License Version 2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions; suitable for both open-source and commercial projects provided you include license attribution.
Quickstart
pip install "misaki[en]"
from misaki import en
g2p = en.G2P(trf=False, british=False, fallback=None)
text = 'Text to convert to phonemes.'
phonemes, tokens = g2p(text)
print(phonemes)
Verify before relying
- Performance characteristics (latency, throughput) for typical text lengths and batch sizes.
- Accuracy metrics or benchmarks against reference datasets.
- Whether transformer-based mode (trf=True) requires additional dependencies or model downloads.
- Compatibility and behavior with out-of-vocabulary or mixed-language text.
- Current state of language-specific features (pitch accent for Japanese, etc.) relative to the description.
Package facts
| License | permissive license permissive |
| Python support | Capped below the current Python release <3.13,>=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesaddictregex |
| Maintenance | Aging 496 days since the last release |
| First released | |
| Downloads | 688,566 / month, #5,340 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: misaki-0.9.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “g2p text to speech”
- misakiConverts written text to phonetic representations…
- g2p-enConverts English text to phoneme sequences using dictionary lookup,…
- phonemizer-forkConverts text to phonetic representations (phones) in multiple…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also g2p-en · g2pkk · pyopenjtalk · sea-g2p · orthography2ipa · kokoro-onnx · phonemizer-fork · ko-speech-tools · phonemizer · kokoro