Subcategories
- Speech (22)
- Analysis (10)
- Conversion (10)
- Players (10)
- MIDI (6)
- Capture/Recording (4)
- Editors (4)
Packages
-
soundfile
Reads and writes audio files in formats like…
permissive · active · 30.0M/mo
-
pydub
Pydub provides a high-level Python interface…
permissive · abandoned · 21.0M/mo
-
mutagen
Mutagen reads and writes audio metadata (tags)…
copyleft · active · 15.0M/mo
-
torchaudio
Provides PyTorch-based audio processing,…
permissive · active · 12.3M/mo
-
moviepy
MoviePy is a Python library for video editing…
permissive · aging · 8.2M/mo
-
tinytag
Reads metadata (artist, title, duration,…
permissive · active · 5.5M/mo
-
livekit-api
Generate access tokens and call LiveKit server…
permissive · active · 5.0M/mo
-
livekit-agents
A framework for building realtime multimodal…
permissive · active · 4.5M/mo
-
pedalboard
pedalboard reads, writes, and processes audio…
copyleft · active · 4.4M/mo
-
ffmpy
ffmpy is a Python wrapper around FFmpeg that…
permissive · aging · 4.3M/mo
-
livekit
Python SDK for building real-time video, audio,…
permissive · active · 4.3M/mo
-
pyloudnorm
Measures and normalizes audio loudness…
permissive · aging · 4.0M/mo
-
livekit-plugins-silero
Provides Silero voice activity detection…
permissive · active · 3.6M/mo
-
livekit-plugins-openai
Integrates OpenAI's Realtime, Responses, LLM,…
permissive · active · 3.5M/mo
-
julius
Julius provides differentiable, GPU-accelerated…
permissive · active · 2.7M/mo
-
livekit-blingfire
Provides BlingFire tokenization bindings for…
permissive · active · 2.7M/mo
-
livekit-plugins-deepgram
Integrates Deepgram's voice AI services…
permissive · active · 2.7M/mo
-
livekit-plugins-turn-detector
Detects end-of-turn in voice conversations for…
permissive · active · 2.5M/mo
-
cloudinary
Cloudinary Python SDK provides image and video…
permissive · active · 2.3M/mo
-
torch-audiomentations
Provides PyTorch-native audio data augmentation…
permissive · aging · 2.1M/mo
-
livekit-plugins-elevenlabs
Integrates ElevenLabs text-to-speech into the…
permissive · active · 1.9M/mo
-
pytgcalls
Enables Python bots to join Telegram voice…
copyleft · abandoned · 1.4M/mo
-
PyAudio
PyAudio provides Python bindings for PortAudio…
permissive · dormant · 1.3M/mo
-
livekit-plugins-cartesia
Integrates Cartesia's voice AI services…
permissive · active · 1.3M/mo
-
pipecat-ai
Pipecat is a Python framework for building…
permissive · active · 1.2M/mo
-
livekit-plugins-google
Integrates Google Cloud AI services (Gemini,…
permissive · active · 1.2M/mo
-
livekit-plugins-anthropic
Integrates Anthropic's Claude models into…
permissive · active · 984.4K/mo
-
pymp4
Parses and builds MP4 box structures from…
permissive · dormant · 931.6K/mo
-
livekit-plugins-assemblyai
Integrates AssemblyAI speech-to-text into the…
permissive · active · 775.7K/mo
-
miniaudio
Miniaudio provides Python bindings for…
permissive · active · 674.9K/mo
-
descript-audio-codec
Compresses audio into discrete codes at 8 kbps…
permissive · active · 487.3K/mo
-
descript-audiotools
Provides object-oriented audio signal handling…
permissive · abandoned · 464.6K/mo
-
livekit-plugins-groq
Integrates Groq's fast inference API with…
permissive · active · 443.5K/mo
-
livekit-plugins-aws
Integrates Amazon AWS AI services (Bedrock,…
permissive · active · 424.8K/mo
-
music21
music21 is a Python toolkit for analyzing,…
permissive · active · 409.5K/mo
-
demucs
Demucs separates music into individual…
permissive · active · 400.5K/mo
-
audio-separator
Separates audio files into multiple stems…
permissive · active · 363.6K/mo
-
encodec
EnCodec is a neural audio codec that compresses…
noncommercial · dormant · 322.8K/mo
-
livekit-plugins-azure
Integrates Azure AI services, particularly…
permissive · active · 311.2K/mo
-
daily-python
A Python SDK for building video and audio…
permissive · active · 296.3K/mo
-
spotdl
spotDL downloads songs from Spotify playlists…
permissive · active · 290.3K/mo
-
samplerate
Wraps libsamplerate (Secret Rabbit Code) to…
permissive · active · 289.1K/mo
-
livekit-plugins-xai
Integrates xAI's Grok LLM with LiveKit's Agent…
permissive · active · 269.4K/mo
-
resemble-perth
Embeds imperceptible watermarks into audio…
permissive · aging · 239.9K/mo
-
livekit-plugins-soniox
Integrates Soniox speech-to-text and…
permissive · active · 229.0K/mo
-
streamlink
Streamlink extracts video streams from various…
permissive · active · 218.0K/mo
-
aubio
aubio is a Python wrapper around a C library…
copyleft · abandoned · 213.3K/mo
-
python-vlc
Python ctypes-based bindings to libvlc that let…
copyleft · aging · 211.7K/mo
-
audiomentations
Audiomentations applies randomized audio…
permissive · active · 211.3K/mo
-
numpy-rms
Calculates Root Mean Square (RMS) values over…
permissive · active · 206.1K/mo
-
numpy-minmax
Finds the minimum and maximum values in a NumPy…
permissive · active · 204.1K/mo
-
coqui-tts
Coqui TTS synthesizes speech from text using…
copyleft · active · 183.7K/mo
-
python-stretch
Pitch-shifts and time-stretches audio using the…
permissive · active · 183.3K/mo
-
pylast
pylast provides a Python interface to Last.fm…
permissive · active · 182.7K/mo
-
livekit-plugins-rime
Integrates Rime speech recognition into…
permissive · active · 177.9K/mo
-
livekit-plugins-speechmatics
Integrates Speechmatics speech-to-text into…
permissive · active · 175.2K/mo
-
pulsectl
Provides a high-level Python interface to…
permissive · dormant · 164.9K/mo
-
soco
SoCo is a Python library for programmatic…
permissive · active · 164.2K/mo
-
pipecat-ai-flows
Manages conversation flows and state machines…
permissive · abandoned · 142.9K/mo
-
silero
Silero provides pre-trained text-to-speech…
permissive · active · 138.2K/mo
-
mediafile
MediaFile reads and writes metadata tags across…
permissive · active · 134.4K/mo
-
pipecat-ai-whisker
Whisker is a real-time debugger for Pipecat…
permissive · active · 112.0K/mo
-
silk-python
Encodes and decodes audio in SILK format, a…
permissive · active · 111.7K/mo
-
TTS
TTS is a deep learning library for…
copyleft · dormant · 108.1K/mo
-
livekit-plugins-telnyx
Integrates Telnyx telephony services with…
permissive · active · 105.8K/mo
-
pipecat-ai-prebuilt
Provides a lightweight, ready-to-use web UI for…
permissive · active · 102.0K/mo
-
sonos-websocket
Async Python library for communicating with…
permissive · active · 101.5K/mo
-
syncedlyrics
Fetches synchronized lyrics in LRC format for…
permissive · dormant · 101.4K/mo
-
python-osc
python-osc implements Open Sound Control (OSC)…
unclear · active · 97.3K/mo
-
typed-ffmpeg
A Python wrapper around FFmpeg with full type…
permissive · active · 93.6K/mo
-
livekit-plugins-gladia
Integrates Gladia's speech-to-text API with…
permissive · active · 93.2K/mo
-
livekit-plugins-sarvam
Integrates Sarvam.ai's Indian-language voice AI…
permissive · active · 91.8K/mo
-
audiofile
Reads and writes audio files across formats…
permissive · active · 90.6K/mo
-
pipecat-ai-small-webrtc-prebuilt
Provides a ready-to-use WebRTC client UI for…
permissive · active · 88.9K/mo
-
livekit-plugins-inworld
Integrates Inworld's text-to-speech and…
permissive · active · 81.6K/mo
-
stftpitchshift
Shifts the pitch and timbre of audio signals…
permissive · aging · 80.8K/mo
-
livekit-plugins-anam
Anam plugin for the LiveKit Agents framework…
permissive · active · 75.0K/mo