$npx skillfedfor your agent

g2p-en

A Simple Python Module for English Grapheme To Phoneme Conversion

With conditionsPyPI LinguisticReleased Dec 20191.1M downloads / moApache Software LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — g2p_en-2.1.0-py3-none-any.whl
v2.1.0 · released 2019-12-31 · 4 runtime deps: numpy, nltk, inflect, distance

Yes, if you need English grapheme-to-phoneme conversion and can tolerate an abandoned package. The module is stable, has no known vulnerabilities, and works well for its narrow purpose. Install it only if you are comfortable maintaining compatibility with its dependencies yourself and do not expect upstream fixes or feature updates. For active projects requiring long-term support, consider whether a maintained alternative exists.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • nltk data files (averaged_perceptron_tagger, cmudict) are downloaded automatically on first run; no explicit setup required but network access is needed during initialization.
  • Low install friction with a pure-Python wheel and four straightforward dependencies.
  • However, the package is abandoned—last release was 2019-12-31 and last commit 2023-01-05—so expect no maintenance, bug fixes, or updates to handle modern dependency versions.

License · maintenance · safety

Apache Software License (permissive) — Apache Software License (permissive) allows free use, modification, and distribution with minimal restrictions, making it safe to integrate into most projects without licensing concerns.

last release 2019-12-31 (2418 days) · last repo commit 2023-01-05 · 931 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,134,313 downloads/mo, #4,313 on PyPI

Verify before relying

pip install g2p_en

from g2p_en import G2p

g2p = G2p()
output = g2p("I have $250 in my pocket.")
print(output)
  • Whether the package's neural model remains accurate for modern English or newly coined words beyond its training data.
  • Compatibility with current versions of numpy, nltk, inflect, and distance given the package's abandoned status.
  • Performance characteristics (latency, memory) for batch processing or real-time speech synthesis pipelines.
Same gist for agents: .md · .json

What it is and what it does

g2p_en is a Python module that converts English text to phoneme sequences, a critical preprocessing step for speech synthesis and text-to-speech systems. It handles the irregularity of English pronunciation by combining three strategies: dictionary lookup via the CMU Pronouncing Dictionary, part-of-speech-based disambiguation for homographs (words with multiple pronunciations), and neural sequence-to-sequence prediction for out-of-vocabulary words. The module also normalizes numbers and abbreviations into spelled-out forms before conversion.

The package depends on numpy for inference (TensorFlow was removed to avoid GPU requirements and API churn), nltk for part-of-speech tagging and dictionary access, inflect for morphological operations, and distance for string matching. It is designed for Python 3.x and installs as a pure wheel with low friction. However, the project has been abandoned since late 2019, with no maintenance or updates since early 2023, so users should be prepared to handle any compatibility issues with modern dependency versions independently.

Use it for

  • Preprocessing text for text-to-speech (TTS) systems where accurate phoneme sequences are required for synthesis quality.
  • Building speech recognition training data pipelines that need ground-truth phoneme alignments from orthographic text.
  • Linguistic analysis of English pronunciation patterns, homograph disambiguation, and out-of-vocabulary word handling.
  • Augmenting speech datasets with phonetic transcriptions for acoustic model training.
  • Educational tools that teach English pronunciation by mapping written words to their phonetic representations.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need English grapheme-to-phoneme conversion and can tolerate an abandoned package.

The module is stable, has no known vulnerabilities, and works well for its narrow purpose. Install it only if you are comfortable maintaining compatibility with its dependencies yourself and do not expect upstream fixes or feature updates. For active projects requiring long-term support, consider whether a maintained alternative exists.

Install

g2p-en on PyPI

Before you install

Low install friction with a pure-Python wheel and four straightforward dependencies. However, the package is abandoned—last release was 2019-12-31 and last commit 2023-01-05—so expect no maintenance, bug fixes, or updates to handle modern dependency versions.

nltk data files (averaged_perceptron_tagger, cmudict) are downloaded automatically on first run; no explicit setup required but network access is needed during initialization.

License in practice

Apache Software License (permissive) allows free use, modification, and distribution with minimal restrictions, making it safe to integrate into most projects without licensing concerns.

Quickstart

pip install g2p_en

from g2p_en import G2p

g2p = G2p()
output = g2p("I have $250 in my pocket.")
print(output)

Verify before relying

  • Whether the package's neural model remains accurate for modern English or newly coined words beyond its training data.
  • Compatibility with current versions of numpy, nltk, inflect, and distance given the package's abandoned status.
  • Performance characteristics (latency, memory) for batch processing or real-time speech synthesis pipelines.

Package facts

LicenseApache Software License permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
numpynltkinflectdistance
MaintenanceAbandoned 2,418 days since the last release
Last repo commit
First released
Downloads1,134,313 / month, #4,313 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: g2p_en-2.1.0-py3-none-any.whl

Tags

Capabilities
grapheme to phoneme conversionenglish text to pronunciationg2p englishphoneme predictionspeech synthesis text preprocessingenglish pronunciation lookuptext to ipa conversion
Topics
speech-synthesisnlpphonetics
PyPI keywords
g2pg2p_eng2pE

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “grapheme to phoneme conversion”

  • g2p-enConverts English text to phoneme sequences using dictionary lookup,…
  • misakiConverts written text to phonetic representations…
  • gruutGruut tokenizes, cleans, and converts text to IPA phonemes for…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also misaki · orthography2ipa · sea-g2p · pronouncing · g2pkk · pyopenjtalk · phonemizer · phonemizer-fork · ko-speech-tools · gruut-lang-en