g2p-en
A Simple Python Module for English Grapheme To Phoneme Conversion
Decision gist · record as of 2026-08-14
Yes, if you need English grapheme-to-phoneme conversion and can tolerate an abandoned package. The module is stable, has no known vulnerabilities, and works well for its narrow purpose. Install it only if you are comfortable maintaining compatibility with its dependencies yourself and do not expect upstream fixes or feature updates. For active projects requiring long-term support, consider whether a maintained alternative exists.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- nltk data files (averaged_perceptron_tagger, cmudict) are downloaded automatically on first run; no explicit setup required but network access is needed during initialization.
- Low install friction with a pure-Python wheel and four straightforward dependencies.
- However, the package is abandoned—last release was 2019-12-31 and last commit 2023-01-05—so expect no maintenance, bug fixes, or updates to handle modern dependency versions.
License · maintenance · safety
Apache Software License (permissive) — Apache Software License (permissive) allows free use, modification, and distribution with minimal restrictions, making it safe to integrate into most projects without licensing concerns.
last release 2019-12-31 (2418 days) · last repo commit 2023-01-05 · 931 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,134,313 downloads/mo, #4,313 on PyPI
Alternatives
Verify before relying
pip install g2p_en
from g2p_en import G2p
g2p = G2p()
output = g2p("I have $250 in my pocket.")
print(output)- Whether the package's neural model remains accurate for modern English or newly coined words beyond its training data.
- Compatibility with current versions of numpy, nltk, inflect, and distance given the package's abandoned status.
- Performance characteristics (latency, memory) for batch processing or real-time speech synthesis pipelines.
What it is and what it does
g2p_en is a Python module that converts English text to phoneme sequences, a critical preprocessing step for speech synthesis and text-to-speech systems. It handles the irregularity of English pronunciation by combining three strategies: dictionary lookup via the CMU Pronouncing Dictionary, part-of-speech-based disambiguation for homographs (words with multiple pronunciations), and neural sequence-to-sequence prediction for out-of-vocabulary words. The module also normalizes numbers and abbreviations into spelled-out forms before conversion.
The package depends on numpy for inference (TensorFlow was removed to avoid GPU requirements and API churn), nltk for part-of-speech tagging and dictionary access, inflect for morphological operations, and distance for string matching. It is designed for Python 3.x and installs as a pure wheel with low friction. However, the project has been abandoned since late 2019, with no maintenance or updates since early 2023, so users should be prepared to handle any compatibility issues with modern dependency versions independently.
Use it for
- Preprocessing text for text-to-speech (TTS) systems where accurate phoneme sequences are required for synthesis quality.
- Building speech recognition training data pipelines that need ground-truth phoneme alignments from orthographic text.
- Linguistic analysis of English pronunciation patterns, homograph disambiguation, and out-of-vocabulary word handling.
- Augmenting speech datasets with phonetic transcriptions for acoustic model training.
- Educational tools that teach English pronunciation by mapping written words to their phonetic representations.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need English grapheme-to-phoneme conversion and can tolerate an abandoned package.
The module is stable, has no known vulnerabilities, and works well for its narrow purpose. Install it only if you are comfortable maintaining compatibility with its dependencies yourself and do not expect upstream fixes or feature updates. For active projects requiring long-term support, consider whether a maintained alternative exists.
Install
g2p-en on PyPI
Before you install
Low install friction with a pure-Python wheel and four straightforward dependencies. However, the package is abandoned—last release was 2019-12-31 and last commit 2023-01-05—so expect no maintenance, bug fixes, or updates to handle modern dependency versions.
nltk data files (averaged_perceptron_tagger, cmudict) are downloaded automatically on first run; no explicit setup required but network access is needed during initialization.
License in practice
Apache Software License (permissive) allows free use, modification, and distribution with minimal restrictions, making it safe to integrate into most projects without licensing concerns.
Quickstart
pip install g2p_en
from g2p_en import G2p
g2p = G2p()
output = g2p("I have $250 in my pocket.")
print(output)
Verify before relying
- Whether the package's neural model remains accurate for modern English or newly coined words beyond its training data.
- Compatibility with current versions of numpy, nltk, inflect, and distance given the package's abandoned status.
- Performance characteristics (latency, memory) for batch processing or real-time speech synthesis pipelines.
Package facts
| License | Apache Software License permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesnumpynltkinflectdistance |
| Maintenance | Abandoned 2,418 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,134,313 / month, #4,313 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: g2p_en-2.1.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “grapheme to phoneme conversion”
- g2p-enConverts English text to phoneme sequences using dictionary lookup,…
- misakiConverts written text to phonetic representations…
- gruutGruut tokenizes, cleans, and converts text to IPA phonemes for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also misaki · orthography2ipa · sea-g2p · pronouncing · g2pkk · pyopenjtalk · phonemizer · phonemizer-fork · ko-speech-tools · gruut-lang-en