$npx skillfedfor your agent

pocketsphinx

Official Python bindings for PocketSphinx

With conditionsPyPI SpeechReleased Jun 2026382.6K downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — pocketsphinx-5.1.1-cp311-cp311-macosx_10_9_universal2.whl · pocketsphinx-5.1.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl · pocketsphinx-5.1.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl
v5.1.1 · released 2026-06-06 · 1 runtime deps: sounddevice

Yes, if you need offline speech recognition without cloud dependencies and can accept that the engine is mature but no longer state-of-the-art. Install friction is moderate (system library required on Linux/macOS), but prebuilt wheels ease setup. No security vulnerabilities are known. Best suited for keyword spotting, voice commands, and prototyping rather than high-accuracy transcription tasks.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • PortAudio library must be installed on the system (libportaudio2 on Debian-like systems) for LiveSpeech to work; AudioFile class does not require it.
  • Medium install friction due to compiled wheels and a system dependency: PortAudio (libportaudio2 on Debian) is required for the LiveSpeech class to function.
  • Prebuilt wheels are available for recent Python versions on common platforms.

License · maintenance · safety

permissive license (permissive) — Licensed under a permissive BSD-like license, which allows commercial and private use with minimal restrictions. No licensing concerns for typical adoption.

last release 2026-06-06 (69 days) · last repo commit 2026-08-10 · 4,332 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 382,553 downloads/mo, #7,086 on PyPI

Verify before relying

pip install pocketsphinx

from pocketsphinx import LiveSpeech
for phrase in LiveSpeech():
    print(phrase)
  • Whether the package's accuracy and model quality meet requirements for production speech recognition tasks.
  • Performance characteristics and latency for real-time transcription on typical hardware.
  • Availability and quality of language models beyond the default US English model.
Same gist for agents: .md · .json

What it is and what it does

PocketSphinx is a mature, speaker-independent continuous speech recognition engine originally developed at Carnegie Mellon University. It provides Python bindings to recognize speech from live microphone streams or pre-recorded audio files, supporting both general speech-to-text transcription and keyword spotting. The package includes built-in language models and dictionaries for US English and can be extended with custom models.

The library offers two main interfaces: LiveSpeech for real-time microphone input and AudioFile for processing recorded audio. It depends on sounddevice for audio I/O and requires the PortAudio library on most systems. Development has largely ceased and the engine is acknowledged to be far from state-of-the-art, but it remains actively maintained and is used in production by many projects.

Use it for

  • Build offline speech-to-text applications that don't require cloud API calls or internet connectivity.
  • Implement keyword spotting or wake-word detection in IoT or embedded voice applications.
  • Process batch audio files to extract transcriptions with custom language models.
  • Create voice command interfaces for desktop or server applications with low latency requirements.
  • Prototype speech recognition features without dependency on external speech recognition services.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need offline speech recognition without cloud dependencies and can accept that the engine is mature but no longer state-of-the-art.

Install friction is moderate (system library required on Linux/macOS), but prebuilt wheels ease setup. No security vulnerabilities are known. Best suited for keyword spotting, voice commands, and prototyping rather than high-accuracy transcription tasks.

Install

pocketsphinx on PyPI

Before you install

Medium install friction due to compiled wheels and a system dependency: PortAudio (libportaudio2 on Debian) is required for the LiveSpeech class to function. Prebuilt wheels are available for recent Python versions on common platforms. The package is actively maintained with a recent release (69 days old) and 4332 repository stars.

PortAudio library must be installed on the system (libportaudio2 on Debian-like systems) for LiveSpeech to work; AudioFile class does not require it.

License in practice

Licensed under a permissive BSD-like license, which allows commercial and private use with minimal restrictions. No licensing concerns for typical adoption.

Quickstart

pip install pocketsphinx

from pocketsphinx import LiveSpeech
for phrase in LiveSpeech():
    print(phrase)

Verify before relying

  • Whether the package's accuracy and model quality meet requirements for production speech recognition tasks.
  • Performance characteristics and latency for real-time transcription on typical hardware.
  • Availability and quality of language models beyond the default US English model.

Package facts

Licensepermissive license permissive
Python supportNot specified
Install frictionMedium. Platform-specific wheel
Runtime dependencies
1 package
sounddevice
MaintenanceActively maintained 69 days since the last release
Last repo commit
First released
Downloads382,553 / month, #7,086 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 6 - MatureLicense :: OSI Approved :: BSD LicenseOperating System :: OS IndependentProgramming Language :: CProgramming Language :: CythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Multimedia :: Sound/Audio :: Speech

Evidence: pocketsphinx-5.1.1-cp311-cp311-macosx_10_9_universal2.whl; pocketsphinx-5.1.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; pocketsphinx-5.1.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; pocketsphinx-5.1.1-cp311-cp311-musllinux_1_2_aarch64.whl; pocketsphinx-5.1.1-cp311-cp311-musllinux_1_2_x86_64.whl; pocketsphinx-5.1.1-cp311-cp311-win_amd64.whl; pocketsphinx-5.1.1-cp313-cp313-macosx_10_13_universal2.whl; pocketsphinx-5.1.1-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; pocketsphinx-5.1.1-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; pocketsphinx-5.1.1-cp313-cp313-musllinux_1_2_aarch64.whl; pocketsphinx-5.1.1-cp313-cp313-musllinux_1_2_x86_64.whl; pocketsphinx-5.1.1-cp313-cp313-win_amd64.whl

Tags

Capabilities
speech recognition pythonspeech to text offlinekeyword spotting audiolive microphone transcriptionaudio file speech recognitionasr library pythoncontinuous speech recognition
Topics
speech-recognitionoffline-asrkeyword-spotting
PyPI keywords
asrspeech

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “speech recognition python”

  • pocketsphinxPocketSphinx provides Python bindings for Carnegie Mellon…
  • voskVosk provides offline speech recognition for 20+ languages, enabling…
  • SpeechRecognitionPerforms speech recognition and transcription using multiple online…

Give your agent the search over MCP, or paste the wish link into any chat.

More Speech packages

SpeechRecognition Worth it
PyPI · Python Modules · released Jun 2026

Performs speech recognition and transcription using multiple online and offline engines, including Google, OpenAI Whisper, CMU Sphinx, and others.

The main gotcha is that most engines require optional dependencies or API credentials.

BSD-3-Clausepure Python · 3.9+
11.7Mdownloads / mo
gTTS With conditions
PyPI · Libraries · released Nov 2024

gTTS converts text to speech using Google Translate's API, writing MP3 audio to files, file-like objects, or stdout via Python library or command-line tool.

However, be aware that it depends on Google Translate's undocumented API—upstream changes can break it without notice, and it is not a substitute for official Google…

MITpure Python · 3.7+
5.3Mdownloads / mo
lhotse Worth it
PyPI · Python Modules · released Apr 2026

Lhotse prepares multimodal (speech, audio, video, image, text) data for machine learning model training with flexible pipelines, on-the-fly augmentation, and efficient data loading.

Install it if you are building speech, audio, or multimodal training pipelines; skip it if you only need simple audio I/O without data augmentation or complex dataset…

Apache-2.0pure Python · 3.8.0+
1.3Mdownloads / mo
piper-tts With conditions
PyPI · Speech · released Aug 2026

Piper TTS is a local neural text-to-speech engine that converts text to speech using embedded phonemization, with support for multiple languages and voices.

copyleftcompiled wheel · 3.9+
891.3Kdownloads / mo
funasr Worth it
PyPI · Python Modules · released Aug 2026

FunASR is a speech recognition toolkit that transcribes audio offline or via streaming, with integrated voice activity detection, speaker identification, punctuation restoration, and emotion/audio-event tagging across multiple languages and deployment targets.

Install it if you need speaker diarization, emotion detection, streaming support, or self-hosted deployment.

MITpure Python · 3.7.0+
497.9Kdownloads / mo
aic-sdk With conditions
PyPI · Speech · released Aug 2026

Python bindings for ai-coustics audio enhancement, voice activity detection, and analysis SDK, supporting real-time audio processing with numpy arrays.

Apache-2.0compiled wheel · 3.10+
296.4Kdownloads / mo

See also sherpa-onnx · sherpa-onnx-core · onnx-asr · openwakeword · vosk · pyobjc-framework-Speech · deepgram-sdk · qwen-asr · whisperx