$npx skillfedfor your agent

Speech packages

22 packages

Packages

SpeechRecognition Worth it
PyPI · Speech · released Jun 2026

Performs speech recognition and transcription using multiple online and offline engines, including Google, OpenAI Whisper, CMU Sphinx, and others.

The main gotcha is that most engines require optional dependencies or API credentials.

BSD-3-Clausepure Python · 3.9+
11.7Mdownloads / mo
gTTS With conditions
PyPI · Speech · released Nov 2024

gTTS converts text to speech using Google Translate's API, writing MP3 audio to files, file-like objects, or stdout via Python library or command-line tool.

However, be aware that it depends on Google Translate's undocumented API—upstream changes can break it without notice, and it is not a substitute for official Google…

MITpure Python · 3.7+
5.3Mdownloads / mo
lhotse Worth it
PyPI · Speech · released Apr 2026

Lhotse prepares multimodal (speech, audio, video, image, text) data for machine learning model training with flexible pipelines, on-the-fly augmentation, and efficient data loading.

Install it if you are building speech, audio, or multimodal training pipelines; skip it if you only need simple audio I/O without data augmentation or complex dataset…

Apache-2.0pure Python · 3.8.0+
1.3Mdownloads / mo
piper-tts With conditions
PyPI · Speech · released Aug 2026

Piper TTS is a local neural text-to-speech engine that converts text to speech using embedded phonemization, with support for multiple languages and voices.

copyleftcompiled wheel · 3.9+
891.3Kdownloads / mo
funasr Worth it
PyPI · Speech · released Aug 2026

FunASR is a speech recognition toolkit that transcribes audio offline or via streaming, with integrated voice activity detection, speaker identification, punctuation restoration, and emotion/audio-event tagging across multiple languages and deployment targets.

Install it if you need speaker diarization, emotion detection, streaming support, or self-hosted deployment.

MITpure Python · 3.7.0+
497.9Kdownloads / mo
pocketsphinx With conditions
PyPI · Speech · released Jun 2026

PocketSphinx provides Python bindings for Carnegie Mellon University's open-source speech recognition engine, enabling continuous speech-to-text and keyword spotting from live microphone input or audio files.

BSD-3-Clausecompiled wheel
382.6Kdownloads / mo
aic-sdk With conditions
PyPI · Speech · released Aug 2026

Python bindings for ai-coustics audio enhancement, voice activity detection, and analysis SDK, supporting real-time audio processing with numpy arrays.

Apache-2.0compiled wheel · 3.10+
296.4Kdownloads / mo
speechmatics-rt Worth it
PyPI · Speech · released Jun 2026

Async Python client for real-time speech-to-text transcription via the Speechmatics API, supporting both single-stream and multi-channel audio processing over WebSocket.

Install it if you need to integrate Speechmatics' real-time API; skip it if you're using a different speech-to-text provider or don't need real-time streaming.

MITpure Python · 3.9+
265.7Kdownloads / mo
onnx-asr Worth it
PyPI · Speech · released Jul 2026

Automatic Speech Recognition using ONNX models with minimal dependencies, supporting multiple modern ASR architectures and running on CPUs, GPUs, and edge devices.

Install it if you need ASR inference in Python without framework overhead.

MITpure Python · 3.10+
230.3Kdownloads / mo
speechmatics-voice With conditions
PyPI · Speech · released Jan 2026

Python SDK for building real-time voice applications with automatic speech segmentation, turn detection, and speaker management on top of the Speechmatics Real-Time API.

MITpure Python · 3.9+
209.5Kdownloads / mo
omnivoice With conditions
PyPI · Speech · released Jul 2026

OmniVoice generates speech from text in over 600 languages using a diffusion-based model, with support for voice cloning from reference audio and voice design via speaker attributes.

Apache-2.0pure Python · 3.10+
206.4Kdownloads / mo
fish-audio-sdk Worth it
PyPI · Speech · released Mar 2026

Official Python client for the Fish Audio API, providing text-to-speech, speech-to-text, voice cloning, and real-time streaming capabilities with both synchronous and asynchronous interfaces.

Apache-2.0pure Python · 3.9+
192.1Kdownloads / mo
coqui-tts With conditions
PyPI · Speech · released Jan 2026

Coqui TTS synthesizes speech from text using deep learning models, supporting over 1100 languages with pretrained weights and tools for training and fine-tuning custom models.

Install it if you want pretrained models out-of-the-box or plan to fine-tune.

MPL-2.0pure Python
183.7Kdownloads / mo
sea-g2p With conditions
PyPI · Speech · released Aug 2026

Converts text to phonemes for Vietnamese, Thai, and Indonesian with English code-switching support, using a memory-mapped binary dictionary and Rust-based engine for fast batch processing.

Apache-2.0compiled wheel · 3.10+
168.6Kdownloads / mo
pvporcupine Worth it
PyPI · Speech · released Jun 2026

Porcupine is a lightweight wake word detection engine that identifies spoken keywords in audio streams, enabling always-listening voice applications with minimal computational overhead.

Apache-2.0pure Python · 3.9+
160.1Kdownloads / mo
monotonic-alignment-search With conditions
PyPI · Speech · released Oct 2025

Finds the most probable alignment between a text sequence and a speech sequence using monotonic alignment search, with Cython-optimized and NumPy implementations.

However, maintenance is minimal (aging status, 303 days since last release); use it as a stable library component rather than expecting active development or rapid…

MITcompiled wheel · 3.10+aging
131.0Kdownloads / mo
TTS With conditions
PyPI · Speech · released Dec 2023

TTS is a deep learning library for text-to-speech synthesis that generates spoken audio from text using pretrained models across multiple languages, with support for model training and fine-tuning.

MPL-2.0compiled wheeldormant
108.1Kdownloads / mo
ko-speech-tools With conditions
PyPI · Speech · released Oct 2025

Provides Korean language processing tools including Hangul romanization, Jamo character conversion, and grapheme-to-phoneme conversion for speech synthesis.

However, verify that the G2P accuracy meets your speech synthesis requirements and understand that maintenance is uncertain given the recent release and small…

Apache-2.0pure Python · 3.10+aging
103.2Kdownloads / mo
syncedlyrics With conditions
PyPI · Speech · released Jul 2024

Fetches synchronized lyrics in LRC format for music tracks from multiple online providers, with options for plain text or word-level karaoke formats.

However, be aware that provider breakage is likely without active maintenance—test each provider you rely on before deploying to production.

MITpure Python · 3.8+dormant
101.4Kdownloads / mo
deepfilternet With conditions
PyPI · Speech · released Aug 2023

DeepFilterNet removes background noise from full-band audio at 48kHz using deep learning models, providing both command-line and Python API interfaces for speech enhancement.

MITpure Pythondormant
78.6Kdownloads / mo
kugelaudio With conditions
PyPI · Speech · released Aug 2026

Official Python SDK for KugelAudio's text-to-speech API, with optional local CPU-based turn detection for conversational applications.

MITpure Python · 3.9+
78.6Kdownloads / mo
speechmatics-batch Worth it
PyPI · Speech · released Jun 2026

Async Python client for submitting audio files to Speechmatics Batch API, monitoring transcription jobs, and retrieving results in multiple formats with support for speaker diarization, translation, and summarization.

MITpure Python · 3.9+
75.4Kdownloads / mo