$npx skillfedfor your agent

Sound/Audio packages

85 packages · page 1 of 3 · 137 including subcategories

Subcategories

Packages

soundfile Worth it
PyPI · Sound/Audio · released Jun 2026

Reads and writes audio files in formats like WAV, FLAC, OGG, and MAT through libsndfile, exposing audio data as NumPy arrays.

BSD-3-Clausepure Python · 3.10+
30.0Mdownloads / mo
pydub With conditions
PyPI · Sound/Audio · released Mar 2021

Pydub provides a high-level Python interface for loading, manipulating, and exporting audio files with simple operations like slicing, concatenation, and format conversion.

However, the abandoned status since 2021-03-10 means no future fixes or compatibility updates—use it only if you can tolerate potential issues with newer Python…

MITpure Pythonabandoned
21.0Mdownloads / mo
mutagen Worth it
PyPI · Sound/Audio · released Jun 2026

Mutagen reads and writes audio metadata (tags) across many audio formats including MP3, FLAC, OGG, MP4, WavPack, and others, supporting ID3v2 and APEv2 tag editing.

Install it if you need to read or edit audio tags programmatically.

GPL-2.0-or-laterpure Python
15.0Mdownloads / mo
torchaudio With conditions
PyPI · Sound/Audio · released Mar 2026

Provides PyTorch-based audio processing, transforms, and dataloaders for machine learning tasks, with GPU acceleration and autograd support for trainable audio features.

BSD-3-Clausecompiled wheel
12.3Mdownloads / mo
moviepy With conditions
PyPI · Sound/Audio · released May 2025

MoviePy is a Python library for video editing that reads, processes, and writes video and audio files by converting them to numpy arrays for frame-level manipulation and effect application.

MITpure Pythonaging
8.2Mdownloads / mo
tinytag Worth it
PyPI · Sound/Audio · released Jul 2026

Reads metadata (artist, title, duration, bitrate, and more) from audio files in formats including MP3, MP4, FLAC, OGG, WAV, and others, without writing or modifying tags.

Install it if you need to extract metadata from audio files without the overhead of a heavier library.

MITpure Python · 3.7+
5.5Mdownloads / mo
livekit-api Worth it
PyPI · Sound/Audio · released Jul 2026

Generate access tokens and call LiveKit server APIs (room management, egress, ingress, SIP, agent dispatch, connectors) from Python backends using async/await.

Install it if you are building a backend that needs to generate tokens, manage rooms, or integrate with LiveKit's services.

Apache-2.0pure Python · 3.9.0+
5.0Mdownloads / mo
livekit-agents Worth it
PyPI · Sound/Audio · released Aug 2026

A framework for building realtime multimodal and voice AI agents that connect to LiveKit rooms and handle audio/video interactions with language models.

Apache-2.0pure Python
4.5Mdownloads / mo
pedalboard With conditions
PyPI · Sound/Audio · released Jul 2026

pedalboard reads, writes, and processes audio files with built-in effects like reverb, distortion, and equalization, plus support for loading VST3 and Audio Unit plugins.

GPL-3.0-onlycompiled wheel · 3.10+
4.4Mdownloads / mo
ffmpy With conditions
PyPI · Sound/Audio · released Nov 2025

ffmpy is a Python wrapper around FFmpeg that lets you build and execute FFmpeg command lines programmatically without writing shell commands directly.

MITpure Python · 3.9+aging
4.3Mdownloads / mo
livekit Worth it
PyPI · Sound/Audio · released Jul 2026

Python SDK for building real-time video, audio, and data applications by connecting to LiveKit servers as a participant or managing rooms via server APIs.

Apache-2.0compiled wheel · 3.9.0+
4.3Mdownloads / mo
pyloudnorm Worth it
PyPI · Sound/Audio · released Jan 2026

Measures and normalizes audio loudness according to the ITU-R BS.1770-4 standard, with support for multiple weighting filters and customizable analysis parameters.

MITpure Python · 3.9+aging
4.0Mdownloads / mo
livekit-plugins-silero Worth it
PyPI · Sound/Audio · released Aug 2026

Provides Silero voice activity detection integration for the LiveKit Agents framework to detect when users are speaking in real-time voice agent applications.

Apache-2.0pure Python · 3.10.0+
3.6Mdownloads / mo
livekit-plugins-openai Worth it
PyPI · Sound/Audio · released Aug 2026

Integrates OpenAI's Realtime, Responses, LLM, TTS, and STT APIs into LiveKit Agents, plus support for OpenAI-compatible providers like Azure OpenAI, Cerebras, Fireworks, Perplexity, and others.

Install it if you're already using LiveKit Agents and want OpenAI integration.

Apache-2.0pure Python · 3.10.0+
3.5Mdownloads / mo
julius Worth it
PyPI · Sound/Audio · released Jun 2026

Julius provides differentiable, GPU-accelerated digital signal processing for audio and 1D signals using PyTorch, including resampling, FFT convolutions, and frequency-domain filtering.

Install it if you need signal processing as part of a neural network or GPU pipeline; skip it if you only do offline audio analysis on CPU.

MITpure Python · 3.9.0+
2.7Mdownloads / mo
livekit-blingfire With conditions
PyPI · Sound/Audio · released Dec 2025

Provides BlingFire tokenization bindings for the LiveKit Agents framework, enabling fast text segmentation and linguistic analysis within voice agent applications.

Apache-2.0compiled wheel · 3.9.0+
2.7Mdownloads / mo
livekit-plugins-deepgram Worth it
PyPI · Sound/Audio · released Aug 2026

Integrates Deepgram's voice AI services (speech-to-text and text-to-speech) into LiveKit Agents for real-time audio processing in agent applications.

Install it if you are building LiveKit Agents that need Deepgram's speech-to-text or text-to-speech services.

Apache-2.0pure Python · 3.10.0+
2.7Mdownloads / mo
livekit-plugins-turn-detector Skip
PyPI · Sound/Audio · released Aug 2026

Detects end-of-turn in voice conversations for LiveKit Agents using a language model trained for this task, replacing simpler voice activity detection with more accurate interruption prevention.

Apache-2.0pure Python · 3.10.0+
2.5Mdownloads / mo
cloudinary Worth it
PyPI · Sound/Audio · released Jul 2026

Cloudinary Python SDK provides image and video upload, transformation, optimization, and delivery through Cloudinary's cloud platform, with built-in Django integration and secure URL generation.

Install it if you're using Cloudinary or considering a managed media platform.

MITpure Python
2.3Mdownloads / mo
torch-audiomentations With conditions
PyPI · Sound/Audio · released Jan 2025

Provides PyTorch-native audio data augmentation transforms that run on CPU or GPU, designed to integrate directly into neural network models as differentiable modules.

However, be aware that maintenance is aging (last release 576 days ago), multiprocessing and multi-GPU setups have known limitations, and some transforms have edge…

MITpure Python · 3.6+aging
2.1Mdownloads / mo
livekit-plugins-elevenlabs With conditions
PyPI · Sound/Audio · released Aug 2026

Integrates ElevenLabs text-to-speech into the LiveKit Agents framework for building realtime voice agents that can speak with ElevenLabs' voices.

Apache-2.0pure Python · 3.10.0+
1.9Mdownloads / mo
pytgcalls Skip
PyPI · Sound/Audio · released Aug 2021

Enables Python bots to join Telegram voice chats, make and receive private calls, and broadcast audio using WebRTC via the tgcalls C++ binding and MTProto protocol.

LGPL-3.0-onlypure Pythonabandoned
1.4Mdownloads / mo
PyAudio With conditions
PyPI · Sound/Audio · released Nov 2023

PyAudio provides Python bindings for PortAudio v19, enabling you to play and record audio on Windows, macOS, and GNU/Linux from Python code.

MITcompiled wheeldormant
1.3Mdownloads / mo
livekit-plugins-cartesia Worth it
PyPI · Sound/Audio · released Aug 2026

Integrates Cartesia's voice AI services (speech-to-text and text-to-speech) into LiveKit Agents for real-time voice applications.

Apache-2.0pure Python · 3.10.0+
1.3Mdownloads / mo
pipecat-ai Worth it
PyPI · Sound/Audio · released Aug 2026

Pipecat is a Python framework for building real-time voice and multimodal conversational AI agents, with support for orchestrating audio, video, AI services, and multi-agent coordination over shared buses or distributed systems.

BSD-2-Clausepure Python · 3.11+
1.2Mdownloads / mo
livekit-plugins-google With conditions
PyPI · Sound/Audio · released Aug 2026

Integrates Google Cloud AI services (Gemini, Speech-to-Text, Text-to-Speech) with LiveKit Agents for building real-time conversational applications.

Apache-2.0pure Python · 3.10+
1.2Mdownloads / mo
opentok With conditions
PyPI · Sound/Audio · released Jun 2026

Server-side SDK for generating OpenTok sessions, tokens, and managing session archives via the Tokbox/OpenTok platform API.

However, if you are starting a new project, verify whether you should use the newer Vonage Server SDK for Python instead, which supports the unified Vonage Video API…

MITpure Python
1.0Mdownloads / mo
livekit-plugins-anthropic With conditions
PyPI · Sound/Audio · released Aug 2026

Integrates Anthropic's Claude models into LiveKit's agent framework, enabling voice agents to use Claude for natural language understanding and generation in real-time conversations.

Apache-2.0pure Python · 3.10.0+
984.4Kdownloads / mo
pymp4 With conditions
PyPI · Sound/Audio · released May 2023

Parses and builds MP4 box structures from binary data using the construct library, enabling programmatic inspection and manipulation of MP4 file metadata.

Install only if you are comfortable with no active bug fixes or feature development; for production use, verify that its scope matches your MP4 variant and box types.

Apache-2.0pure Pythondormant
931.6Kdownloads / mo
mido Worth it
PyPI · Sound/Audio · released Oct 2024

Mido provides a Python library for creating, parsing, and sending MIDI messages and files, with support for multiple backends (RtMidi, PortMidi, Pygame) and port I/O.

Install it if you need to work with MIDI messages, files, or hardware in any capacity.

MITpure Python
893.4Kdownloads / mo
livekit-plugins-assemblyai With conditions
PyPI · Sound/Audio · released Aug 2026

Integrates AssemblyAI speech-to-text into the LiveKit Agents framework for building real-time voice agents.

Apache-2.0pure Python · 3.10.0+
775.7Kdownloads / mo
miniaudio Worth it
PyPI · Sound/Audio · released Apr 2026

Miniaudio provides Python bindings for cross-platform audio playback, recording, decoding, and sample format conversion using the miniaudio C library.

MITcompiled wheel · 3.8+
674.9Kdownloads / mo
descript-audio-codec With conditions
PyPI · Sound/Audio · released Jul 2023

Compresses audio into discrete codes at 8 kbps bitrate and reconstructs it with high fidelity, supporting 16 kHz, 24 kHz, and 44.1 kHz sampling rates across speech, music, and environmental audio.

permissive licensepure Python
487.3Kdownloads / mo
descript-audiotools Skip
PyPI · Sound/Audio · released Jul 2023

Provides object-oriented audio signal handling with fast augmentation, batching, padding, and playback capabilities for audio processing workflows.

permissive licensepure Pythonabandoned
464.6Kdownloads / mo
livekit-plugins-groq Worth it
PyPI · Sound/Audio · released Aug 2026

Integrates Groq's fast inference API with LiveKit Agents, enabling real-time LLM responses for voice and conversational agents.

Install it if you are already using LiveKit Agents and want to use Groq as your LLM provider, or if you are evaluating Groq for real-time voice agent workloads.

Apache-2.0pure Python · 3.10.0+
443.5Kdownloads / mo
livekit-plugins-aws Worth it
PyPI · Sound/Audio · released Aug 2026

Integrates Amazon AWS AI services (Bedrock, Polly, Transcribe, Nova) into LiveKit Agents for speech-to-speech, text-to-speech, speech-to-text, and LLM capabilities.

Install it if you are building LiveKit voice agents and need AWS AI services (Bedrock, Transcribe, Polly, Nova).

Apache-2.0pure Python · 3.10.0+
424.8Kdownloads / mo
music21 With conditions
PyPI · Sound/Audio · released Jun 2026

music21 is a Python toolkit for analyzing, processing, and computationally working with musical scores and MIDI data, supporting symbolic music representation and analysis workflows.

BSD-3-Clausepure Python · 3.11+
409.5Kdownloads / mo
demucs Worth it
PyPI · Sound/Audio · released Jul 2026

Demucs separates music into individual stems—drums, bass, vocals, and accompaniment—using a hybrid transformer-based neural network trained on waveform and spectrogram domains.

MITpure Python · 3.10+
400.5Kdownloads / mo
audio-separator With conditions
PyPI · Sound/Audio · released Jul 2026

Separates audio files into multiple stems (vocals, instruments, drums, bass, etc.) using pre-trained deep learning models, available as a CLI tool or Python library.

The main gotcha is the separate FFmpeg system requirement and the need to choose an appropriate hardware acceleration path (CUDA, CoreML, CPU, or experimental DirectML).

MITpure Python · 3.10+
363.6Kdownloads / mo
encodec With conditions
PyPI · Sound/Audio · released Oct 2022

EnCodec is a neural audio codec that compresses audio to low bitrates while preserving high fidelity, with separate models for 24 kHz mono and 48 kHz stereo audio.

non-commercialbuilds from source · 3.8.0+dormant
322.8Kdownloads / mo