Packages
Performs speech recognition and transcription using multiple online and offline engines, including Google, OpenAI Whisper, CMU Sphinx, and others.
The main gotcha is that most engines require optional dependencies or API credentials.
gTTS converts text to speech using Google Translate's API, writing MP3 audio to files, file-like objects, or stdout via Python library or command-line tool.
However, be aware that it depends on Google Translate's undocumented API—upstream changes can break it without notice, and it is not a substitute for official Google…
Lhotse prepares multimodal (speech, audio, video, image, text) data for machine learning model training with flexible pipelines, on-the-fly augmentation, and efficient data loading.
Install it if you are building speech, audio, or multimodal training pipelines; skip it if you only need simple audio I/O without data augmentation or complex dataset…
Piper TTS is a local neural text-to-speech engine that converts text to speech using embedded phonemization, with support for multiple languages and voices.
FunASR is a speech recognition toolkit that transcribes audio offline or via streaming, with integrated voice activity detection, speaker identification, punctuation restoration, and emotion/audio-event tagging across multiple languages and deployment targets.
Install it if you need speaker diarization, emotion detection, streaming support, or self-hosted deployment.
PocketSphinx provides Python bindings for Carnegie Mellon University's open-source speech recognition engine, enabling continuous speech-to-text and keyword spotting from live microphone input or audio files.
Python bindings for ai-coustics audio enhancement, voice activity detection, and analysis SDK, supporting real-time audio processing with numpy arrays.
Async Python client for real-time speech-to-text transcription via the Speechmatics API, supporting both single-stream and multi-channel audio processing over WebSocket.
Install it if you need to integrate Speechmatics' real-time API; skip it if you're using a different speech-to-text provider or don't need real-time streaming.
Automatic Speech Recognition using ONNX models with minimal dependencies, supporting multiple modern ASR architectures and running on CPUs, GPUs, and edge devices.
Install it if you need ASR inference in Python without framework overhead.
Python SDK for building real-time voice applications with automatic speech segmentation, turn detection, and speaker management on top of the Speechmatics Real-Time API.
OmniVoice generates speech from text in over 600 languages using a diffusion-based model, with support for voice cloning from reference audio and voice design via speaker attributes.
Official Python client for the Fish Audio API, providing text-to-speech, speech-to-text, voice cloning, and real-time streaming capabilities with both synchronous and asynchronous interfaces.
Coqui TTS synthesizes speech from text using deep learning models, supporting over 1100 languages with pretrained weights and tools for training and fine-tuning custom models.
Install it if you want pretrained models out-of-the-box or plan to fine-tune.
Converts text to phonemes for Vietnamese, Thai, and Indonesian with English code-switching support, using a memory-mapped binary dictionary and Rust-based engine for fast batch processing.
Porcupine is a lightweight wake word detection engine that identifies spoken keywords in audio streams, enabling always-listening voice applications with minimal computational overhead.
Finds the most probable alignment between a text sequence and a speech sequence using monotonic alignment search, with Cython-optimized and NumPy implementations.
However, maintenance is minimal (aging status, 303 days since last release); use it as a stable library component rather than expecting active development or rapid…
TTS is a deep learning library for text-to-speech synthesis that generates spoken audio from text using pretrained models across multiple languages, with support for model training and fine-tuning.
Provides Korean language processing tools including Hangul romanization, Jamo character conversion, and grapheme-to-phoneme conversion for speech synthesis.
However, verify that the G2P accuracy meets your speech synthesis requirements and understand that maintenance is uncertain given the recent release and small…
Fetches synchronized lyrics in LRC format for music tracks from multiple online providers, with options for plain text or word-level karaoke formats.
However, be aware that provider breakage is likely without active maintenance—test each provider you rely on before deploying to production.
DeepFilterNet removes background noise from full-band audio at 48kHz using deep learning models, providing both command-line and Python API interfaces for speech enhancement.
Official Python SDK for KugelAudio's text-to-speech API, with optional local CPU-based turn detection for conversational applications.
Async Python client for submitting audio files to Speechmatics Batch API, monitoring transcription jobs, and retrieving results in multiple formats with support for speaker diarization, translation, and summarization.