pyannote-audio
State-of-the-art speaker diarization toolkit
Decision gist · record as of 2026-08-14
Yes, if you need speaker diarization. The package is actively maintained, has no known vulnerabilities, and offers both free and premium options. Install friction is low (pure wheel, standard ML dependencies). The main gotcha is the ffmpeg system dependency and the 21-package dependency tree—typical for PyTorch-based audio ML but not lightweight. Requires Python >=3.10. Suitable for production use with either the open-source community-1 model or the cloud-hosted precision-2 service.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- ffmpeg must be installed on your system; requires Python >=3.10; Hugging Face access token needed for community-1 model
- Low install friction with a pure-Python wheel.
- Active maintenance (45 days since last release).
License · maintenance · safety
(unclear)
last release 2026-06-30 (45 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,306,394 downloads/mo, #3,152 on PyPI
Alternatives
Verify before relying
pip install pyannote-audio
import torch
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1",
token="HUGGINGFACE_ACCESS_TOKEN")
pipeline.to(torch.device("cuda"))
output = pipeline("audio.wav")
for turn, speaker in output.speaker_diarization:
print(f"start={turn.start:.1f}s stop={turn.end:.1f}s speaker_{speaker}")- Whether the community-1 model requires accepting terms on Hugging Face before first use
- GPU memory requirements for different audio file durations
- Whether telemetry can be fully disabled without environment variable configuration
What it is and what it does
pyannote.audio is a speaker diarization toolkit built on PyTorch that answers the question 'who spoke when' in audio files. It provides pretrained pipelines ready to use out of the box: a free community-1 open-source model that runs locally, and a premium precision-2 model that runs on pyannoteAI servers. Both pipelines take an audio file and output time-segmented speaker labels, identifying speaker boundaries and assigning consistent speaker IDs across the file.
The package is designed for developers and researchers who need to process audio programmatically. It handles audio decoding via torchcodec (requiring ffmpeg), manages model downloads from huggingface-hub, and supports GPU acceleration. The toolkit includes optional telemetry that tracks pipeline usage and file durations in a privacy-preserving way, configurable via environment variable or Python API. It is actively maintained and widely used.
Use it for
- Transcription preprocessing: identify speaker boundaries before passing segments to speech-to-text
- Meeting analysis: determine who spoke when in recorded meetings or conference calls
- Podcast or interview processing: separate and label different speakers for editing or analysis
- Audio quality assurance: detect speaker overlap or anomalies in multi-speaker recordings
- Voice biometrics: extract speaker segments for voiceprinting or speaker verification tasks
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need speaker diarization.
The package is actively maintained, has no known vulnerabilities, and offers both free and premium options. Install friction is low (pure wheel, standard ML dependencies). The main gotcha is the ffmpeg system dependency and the 21-package dependency tree—typical for PyTorch-based audio ML but not lightweight. Requires Python >=3.10. Suitable for production use with either the open-source community-1 model or the cloud-hosted precision-2 service.
Install
pyannote-audio on PyPI
Before you install
Low install friction with a pure-Python wheel. Active maintenance (45 days since last release). Requires ffmpeg as a system dependency for audio decoding, and 21 runtime dependencies including torch, torchaudio, and lightning—a substantial but standard ML stack.
ffmpeg must be installed on your system; requires Python >=3.10; Hugging Face access token needed for community-1 model
Quickstart
pip install pyannote-audio
import torch
from pyannote.audio import Pipeline
pipeline = Pipeline.from_pretrained(
"pyannote/speaker-diarization-community-1",
token="HUGGINGFACE_ACCESS_TOKEN")
pipeline.to(torch.device("cuda"))
output = pipeline("audio.wav")
for turn, speaker in output.speaker_diarization:
print(f"start={turn.start:.1f}s stop={turn.end:.1f}s speaker_{speaker}")
Verify before relying
- Whether the community-1 model requires accepting terms on Hugging Face before first use
- GPU memory requirements for different audio file durations
- Whether telemetry can be fully disabled without environment variable configuration
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 21 packagesasteroid-filterbankseinopshuggingface-hublightningmatplotlibopentelemetry-apiopentelemetry-exporter-otlpopentelemetry-sdkpyannote-corepyannote-databasepyannote-metricspyannote-pipelinepyannoteai-sdkpytorch-metric-learningrichsafetensorstorch-audiomentationstorchtorchaudiotorchcodectorchmetrics |
| Maintenance | Actively maintained 45 days since the last release |
| First released | |
| Downloads | 2,306,394 / month, #3,152 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: pyannote_audio-4.0.7-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “who spoke when audio”
- pyannote-audioSpeaker diarization toolkit that identifies and separates individual…
- pyannoteai-sdkClient library for pyannoteAI's speaker diarization, speaker…
- pyannote-metricsEvaluates and analyzes speaker diarization systems by computing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also pyannote-core · pyannote-database · pyannote-metrics · pyannoteai-sdk · Resemblyzer · speechbrain · sherpa-onnx-core · sherpa-onnx · whisperx · speechmatics-voice