$npx skillfedfor your agent

pyannote-audio

State-of-the-art speaker diarization toolkit

With conditionsPyPI Artificial IntelligenceReleased Jun 20262.3M downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pyannote_audio-4.0.7-py3-none-any.whl
v4.0.7 · released 2026-06-30 · Python >=3.10 · 21 runtime deps: asteroid-filterbanks, einops, huggingface-hub, lightning, matplotlib, opentelemetry-api, opentelemetry-exporter-otlp, opentelemetry-sdk

Yes, if you need speaker diarization. The package is actively maintained, has no known vulnerabilities, and offers both free and premium options. Install friction is low (pure wheel, standard ML dependencies). The main gotcha is the ffmpeg system dependency and the 21-package dependency tree—typical for PyTorch-based audio ML but not lightweight. Requires Python >=3.10. Suitable for production use with either the open-source community-1 model or the cloud-hosted precision-2 service.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • ffmpeg must be installed on your system; requires Python >=3.10; Hugging Face access token needed for community-1 model
  • Low install friction with a pure-Python wheel.
  • Active maintenance (45 days since last release).

License · maintenance · safety

(unclear)

last release 2026-06-30 (45 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,306,394 downloads/mo, #3,152 on PyPI

Verify before relying

pip install pyannote-audio

import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token="HUGGINGFACE_ACCESS_TOKEN")
pipeline.to(torch.device("cuda"))
output = pipeline("audio.wav")

for turn, speaker in output.speaker_diarization:
    print(f"start={turn.start:.1f}s stop={turn.end:.1f}s speaker_{speaker}")
  • Whether the community-1 model requires accepting terms on Hugging Face before first use
  • GPU memory requirements for different audio file durations
  • Whether telemetry can be fully disabled without environment variable configuration
Same gist for agents: .md · .json

What it is and what it does

pyannote.audio is a speaker diarization toolkit built on PyTorch that answers the question 'who spoke when' in audio files. It provides pretrained pipelines ready to use out of the box: a free community-1 open-source model that runs locally, and a premium precision-2 model that runs on pyannoteAI servers. Both pipelines take an audio file and output time-segmented speaker labels, identifying speaker boundaries and assigning consistent speaker IDs across the file.

The package is designed for developers and researchers who need to process audio programmatically. It handles audio decoding via torchcodec (requiring ffmpeg), manages model downloads from huggingface-hub, and supports GPU acceleration. The toolkit includes optional telemetry that tracks pipeline usage and file durations in a privacy-preserving way, configurable via environment variable or Python API. It is actively maintained and widely used.

Use it for

  • Transcription preprocessing: identify speaker boundaries before passing segments to speech-to-text
  • Meeting analysis: determine who spoke when in recorded meetings or conference calls
  • Podcast or interview processing: separate and label different speakers for editing or analysis
  • Audio quality assurance: detect speaker overlap or anomalies in multi-speaker recordings
  • Voice biometrics: extract speaker segments for voiceprinting or speaker verification tasks

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need speaker diarization.

The package is actively maintained, has no known vulnerabilities, and offers both free and premium options. Install friction is low (pure wheel, standard ML dependencies). The main gotcha is the ffmpeg system dependency and the 21-package dependency tree—typical for PyTorch-based audio ML but not lightweight. Requires Python >=3.10. Suitable for production use with either the open-source community-1 model or the cloud-hosted precision-2 service.

Install

pyannote-audio on PyPI

Before you install

Low install friction with a pure-Python wheel. Active maintenance (45 days since last release). Requires ffmpeg as a system dependency for audio decoding, and 21 runtime dependencies including torch, torchaudio, and lightning—a substantial but standard ML stack.

ffmpeg must be installed on your system; requires Python >=3.10; Hugging Face access token needed for community-1 model

Quickstart

pip install pyannote-audio

import torch
from pyannote.audio import Pipeline

pipeline = Pipeline.from_pretrained(
    "pyannote/speaker-diarization-community-1",
    token="HUGGINGFACE_ACCESS_TOKEN")
pipeline.to(torch.device("cuda"))
output = pipeline("audio.wav")

for turn, speaker in output.speaker_diarization:
    print(f"start={turn.start:.1f}s stop={turn.end:.1f}s speaker_{speaker}")

Verify before relying

  • Whether the community-1 model requires accepting terms on Hugging Face before first use
  • GPU memory requirements for different audio file durations
  • Whether telemetry can be fully disabled without environment variable configuration

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
21 packages
asteroid-filterbankseinopshuggingface-hublightningmatplotlibopentelemetry-apiopentelemetry-exporter-otlpopentelemetry-sdkpyannote-corepyannote-databasepyannote-metricspyannote-pipelinepyannoteai-sdkpytorch-metric-learningrichsafetensorstorch-audiomentationstorchtorchaudiotorchcodectorchmetrics
MaintenanceActively maintained 45 days since the last release
First released
Downloads2,306,394 / month, #3,152 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: pyannote_audio-4.0.7-py3-none-any.whl

Tags

Capabilities
speaker diarizationwho spoke when audiospeaker segmentationspeaker identification audiovoice activity detectionspeaker separationaudio speaker tracking
Topics
audio-processingspeaker-diarizationpytorch-ml

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “who spoke when audio”

  • pyannote-audioSpeaker diarization toolkit that identifies and separates individual…
  • pyannoteai-sdkClient library for pyannoteAI's speaker diarization, speaker…
  • pyannote-metricsEvaluates and analyzes speaker diarization systems by computing…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also pyannote-core · pyannote-database · pyannote-metrics · pyannoteai-sdk · Resemblyzer · speechbrain · sherpa-onnx-core · sherpa-onnx · whisperx · speechmatics-voice