Subcategories
Packages
Reads and writes audio files in formats like WAV, FLAC, OGG, and MAT through libsndfile, exposing audio data as NumPy arrays.
Pydub provides a high-level Python interface for loading, manipulating, and exporting audio files with simple operations like slicing, concatenation, and format conversion.
However, the abandoned status since 2021-03-10 means no future fixes or compatibility updates—use it only if you can tolerate potential issues with newer Python…
Mutagen reads and writes audio metadata (tags) across many audio formats including MP3, FLAC, OGG, MP4, WavPack, and others, supporting ID3v2 and APEv2 tag editing.
Install it if you need to read or edit audio tags programmatically.
Provides PyTorch-based audio processing, transforms, and dataloaders for machine learning tasks, with GPU acceleration and autograd support for trainable audio features.
MoviePy is a Python library for video editing that reads, processes, and writes video and audio files by converting them to numpy arrays for frame-level manipulation and effect application.
Reads metadata (artist, title, duration, bitrate, and more) from audio files in formats including MP3, MP4, FLAC, OGG, WAV, and others, without writing or modifying tags.
Install it if you need to extract metadata from audio files without the overhead of a heavier library.
Generate access tokens and call LiveKit server APIs (room management, egress, ingress, SIP, agent dispatch, connectors) from Python backends using async/await.
Install it if you are building a backend that needs to generate tokens, manage rooms, or integrate with LiveKit's services.
A framework for building realtime multimodal and voice AI agents that connect to LiveKit rooms and handle audio/video interactions with language models.
pedalboard reads, writes, and processes audio files with built-in effects like reverb, distortion, and equalization, plus support for loading VST3 and Audio Unit plugins.
ffmpy is a Python wrapper around FFmpeg that lets you build and execute FFmpeg command lines programmatically without writing shell commands directly.
Python SDK for building real-time video, audio, and data applications by connecting to LiveKit servers as a participant or managing rooms via server APIs.
Measures and normalizes audio loudness according to the ITU-R BS.1770-4 standard, with support for multiple weighting filters and customizable analysis parameters.
Provides Silero voice activity detection integration for the LiveKit Agents framework to detect when users are speaking in real-time voice agent applications.
Integrates OpenAI's Realtime, Responses, LLM, TTS, and STT APIs into LiveKit Agents, plus support for OpenAI-compatible providers like Azure OpenAI, Cerebras, Fireworks, Perplexity, and others.
Install it if you're already using LiveKit Agents and want OpenAI integration.
Julius provides differentiable, GPU-accelerated digital signal processing for audio and 1D signals using PyTorch, including resampling, FFT convolutions, and frequency-domain filtering.
Install it if you need signal processing as part of a neural network or GPU pipeline; skip it if you only do offline audio analysis on CPU.
Provides BlingFire tokenization bindings for the LiveKit Agents framework, enabling fast text segmentation and linguistic analysis within voice agent applications.
Integrates Deepgram's voice AI services (speech-to-text and text-to-speech) into LiveKit Agents for real-time audio processing in agent applications.
Install it if you are building LiveKit Agents that need Deepgram's speech-to-text or text-to-speech services.
Detects end-of-turn in voice conversations for LiveKit Agents using a language model trained for this task, replacing simpler voice activity detection with more accurate interruption prevention.
Cloudinary Python SDK provides image and video upload, transformation, optimization, and delivery through Cloudinary's cloud platform, with built-in Django integration and secure URL generation.
Install it if you're using Cloudinary or considering a managed media platform.
Provides PyTorch-native audio data augmentation transforms that run on CPU or GPU, designed to integrate directly into neural network models as differentiable modules.
However, be aware that maintenance is aging (last release 576 days ago), multiprocessing and multi-GPU setups have known limitations, and some transforms have edge…
Integrates ElevenLabs text-to-speech into the LiveKit Agents framework for building realtime voice agents that can speak with ElevenLabs' voices.
Enables Python bots to join Telegram voice chats, make and receive private calls, and broadcast audio using WebRTC via the tgcalls C++ binding and MTProto protocol.
PyAudio provides Python bindings for PortAudio v19, enabling you to play and record audio on Windows, macOS, and GNU/Linux from Python code.
Integrates Cartesia's voice AI services (speech-to-text and text-to-speech) into LiveKit Agents for real-time voice applications.
Pipecat is a Python framework for building real-time voice and multimodal conversational AI agents, with support for orchestrating audio, video, AI services, and multi-agent coordination over shared buses or distributed systems.
Integrates Google Cloud AI services (Gemini, Speech-to-Text, Text-to-Speech) with LiveKit Agents for building real-time conversational applications.
Server-side SDK for generating OpenTok sessions, tokens, and managing session archives via the Tokbox/OpenTok platform API.
However, if you are starting a new project, verify whether you should use the newer Vonage Server SDK for Python instead, which supports the unified Vonage Video API…
Integrates Anthropic's Claude models into LiveKit's agent framework, enabling voice agents to use Claude for natural language understanding and generation in real-time conversations.
Parses and builds MP4 box structures from binary data using the construct library, enabling programmatic inspection and manipulation of MP4 file metadata.
Install only if you are comfortable with no active bug fixes or feature development; for production use, verify that its scope matches your MP4 variant and box types.
Mido provides a Python library for creating, parsing, and sending MIDI messages and files, with support for multiple backends (RtMidi, PortMidi, Pygame) and port I/O.
Install it if you need to work with MIDI messages, files, or hardware in any capacity.
Integrates AssemblyAI speech-to-text into the LiveKit Agents framework for building real-time voice agents.
Miniaudio provides Python bindings for cross-platform audio playback, recording, decoding, and sample format conversion using the miniaudio C library.
Compresses audio into discrete codes at 8 kbps bitrate and reconstructs it with high fidelity, supporting 16 kHz, 24 kHz, and 44.1 kHz sampling rates across speech, music, and environmental audio.
Provides object-oriented audio signal handling with fast augmentation, batching, padding, and playback capabilities for audio processing workflows.
Integrates Groq's fast inference API with LiveKit Agents, enabling real-time LLM responses for voice and conversational agents.
Install it if you are already using LiveKit Agents and want to use Groq as your LLM provider, or if you are evaluating Groq for real-time voice agent workloads.
Integrates Amazon AWS AI services (Bedrock, Polly, Transcribe, Nova) into LiveKit Agents for speech-to-speech, text-to-speech, speech-to-text, and LLM capabilities.
Install it if you are building LiveKit voice agents and need AWS AI services (Bedrock, Transcribe, Polly, Nova).
music21 is a Python toolkit for analyzing, processing, and computationally working with musical scores and MIDI data, supporting symbolic music representation and analysis workflows.
Demucs separates music into individual stems—drums, bass, vocals, and accompaniment—using a hybrid transformer-based neural network trained on waveform and spectrogram domains.
Separates audio files into multiple stems (vocals, instruments, drums, bass, etc.) using pre-trained deep learning models, available as a CLI tool or Python library.
The main gotcha is the separate FFmpeg system requirement and the need to choose an appropriate hardware acceleration path (CUDA, CoreML, CPU, or experimental DirectML).
EnCodec is a neural audio codec that compresses audio to low bitrates while preserving high fidelity, with separate models for 24 kHz mono and 48 kHz stereo audio.