$npx skillfedfor your agent

phonemizer

Simple text to phones converter for multiple languages

Worth itPyPI LinguisticReleased Jul 2026431.8K downloads / mocopyleft licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — phonemizer-3.4.0-py3-none-any.whl
v3.4.0 · released 2026-07-31 · Python >=3.8 · 4 runtime deps: joblib, attrs, dlinfo, typing-extensions

Yes. Phonemizer is a well-maintained, actively developed tool with low install friction and no known vulnerabilities. It fills a clear niche in phonetic text processing across many languages. The GPLv3+ license is a constraint only if you need to build proprietary closed-source software; for research, open-source, and academic use, it is unencumbered. The main gotcha is that you must install one of the external backends separately, but that is by design and well-documented.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires one of the external TTS/phonemization backends (espeak-ng, festival, or segments) to be installed on the system; the Python package alone is a wrapper.
  • Low install friction: pure Python wheel with four lightweight runtime dependencies (joblib, attrs, dlinfo, typing-extensions).
  • Active maintenance with a release 14 days ago and 1567 repository stars.

License · maintenance · safety

copyleft license (copyleft) — GPLv3+ copyleft license requires that any derivative work or distribution must also be licensed under GPLv3 or later. Suitable for open-source projects but incompatible with proprietary closed-source applications.

last release 2026-07-31 (14 days) · last repo commit 2026-08-04 · 1,567 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 431,810 downloads/mo, #6,714 on PyPI

Verify before relying

pip install phonemizer
from phonemizer.phonemize import phonemize
result = phonemize('hello', language='en-us', backend='espeak')
  • Whether backend installation (espeak-ng, festival, segments) is automatic or requires manual system setup on each platform.
  • Performance characteristics and memory usage for large-scale batch phonemization tasks.
  • Accuracy and coverage differences between the four backends for specific language pairs.
Same gist for agents: .md · .json

What it is and what it does

Phonemizer is a Python wrapper around multiple text-to-speech and phoneme-extraction backends that converts written words and sentences into their phonetic representations. It supports four backends—espeak-ng (IPA output, 100+ languages), espeak-mbrola (SAMPA, 35 languages), festival (US English, syllable-level tokenization), and segments (user-defined grapheme-to-phoneme mappings)—each with different speed, language coverage, and output format trade-offs. You choose which backend to use based on your language, required phoneme alphabet, and whether you need syllable-level or word-level boundaries.

The package provides both a command-line tool (`phonemize`) and a Python API (`phonemizer.phonemize`). It is actively maintained, has low install friction (pure Python with minimal dependencies), and is published under GPLv3+. The main constraint is that it requires at least one external backend system library to be installed separately; the package itself is a Python interface to those tools.

Use it for

  • Generate IPA transcriptions for linguistic research or speech corpus annotation across many languages.
  • Prepare phonetic training data for automatic speech recognition (ASR) or text-to-speech (TTS) systems.
  • Extract syllable-level phonetic boundaries for prosody analysis or speech synthesis in festival-supported contexts.
  • Build custom phonemization pipelines using user-defined grapheme-to-phoneme mappings via the segments backend.
  • Batch-process large text corpora into phonetic form for phonological or acoustic studies.
  • Integrate phonetic transcription into NLP pipelines for multilingual phonetic feature extraction.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Phonemizer is a well-maintained, actively developed tool with low install friction and no known vulnerabilities. It fills a clear niche in phonetic text processing across many languages. The GPLv3+ license is a constraint only if you need to build proprietary closed-source software; for research, open-source, and academic use, it is unencumbered. The main gotcha is that you must install one of the external backends separately, but that is by design and well-documented.

Install

phonemizer on PyPI

Before you install

Low install friction: pure Python wheel with four lightweight runtime dependencies (joblib, attrs, dlinfo, typing-extensions). Active maintenance with a release 14 days ago and 1567 repository stars.

Requires one of the external TTS/phonemization backends (espeak-ng, festival, or segments) to be installed on the system; the Python package alone is a wrapper.

License in practice

GPLv3+ copyleft license requires that any derivative work or distribution must also be licensed under GPLv3 or later. Suitable for open-source projects but incompatible with proprietary closed-source applications.

Quickstart

pip install phonemizer
from phonemizer.phonemize import phonemize
result = phonemize('hello', language='en-us', backend='espeak')

Verify before relying

  • Whether backend installation (espeak-ng, festival, segments) is automatic or requires manual system setup on each platform.
  • Performance characteristics and memory usage for large-scale batch phonemization tasks.
  • Accuracy and coverage differences between the four backends for specific language pairs.

Package facts

Licensecopyleft license copyleft
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
joblibattrsdlinfotyping-extensions
MaintenanceActively maintained 14 days since the last release
Last repo commit
First released
Downloads431,810 / month, #6,714 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: GNU General Public License v3 or later (GPLv3+)Operating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: phonemizer-3.4.0-py3-none-any.whl

Tags

Capabilities
text to phonemes conversionIPA phonetic transcriptiongrapheme to phoneme g2pmultilingual phonemizationespeak python wrapperspeech phoneme extractionlinguistic text processing
Topics
phoneticsmultilingualspeech-processing
PyPI keywords
linguisticsG2PphoneespeakfestivalTTS

Let your AI agent find packages like this

Example. Real query, live index.

An agent finds packages by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language. Give your agent the search over MCP.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also espeakng-loader · gruut-ipa · panphon · phonemizer-fork · orthography2ipa · epitran · misaki · g2p-en · gruut-lang-en · gruut