pyannote-database
Interface to multimedia databases and experimental protocols
Decision gist · record as of 2026-08-14
Yes, if you are working with speaker diarization, speaker verification, or other multimedia ML tasks and need a standardized way to define and iterate over dataset splits. The low install friction and stable dependency set make it straightforward to adopt. However, verify the license terms first (they are not declared in the package metadata), and be aware that maintenance is aging—expect no rapid updates but likely sufficient stability for established workflows.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- YAML configuration file must be created and referenced; automatic loading checks ~/.pyannote/database.yml, ./database.yml, and PYANNOTE_DATABASE_CONFIG environment variable.
- Low install friction with a pure-Python wheel and minimal dependencies (pandas, pyannote-core, pyyaml).
License · maintenance · safety
(unclear) — License treatment is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.
last release 2025-12-07 (250 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,966,223 downloads/mo, #2,805 on PyPI
Alternatives
Verify before relying
pip install pyannote.database
from pyannote.database import registry
registry.load_database("/path/to/database.yml")
protocol = registry.get_protocol('MyDatabase.Protocol.MyProtocol')
for resource in protocol.train():
print(resource["uri"])- Whether the package is actively maintained or in long-term stable mode despite the 250-day release gap.
- What license actually governs this package (SPDX and raw license data are both missing).
- Whether custom data loaders and preprocessors cover common audio/video formats beyond RTTM, UEM, and CTM.
What it is and what it does
pyannote-database is a framework for defining and iterating over multimedia datasets with reproducible experimental protocols. It models resources (audio files, video files, images, etc.) as protocol files with URIs and associated metadata, then organizes them into train, development, and test subsets via YAML configuration. The package handles lazy loading and caching of metadata through pluggable data loaders (built-in support for RTTM, UEM, and CTM formats) and allows on-the-fly augmentation via preprocessors.
Typically used in speech processing and speaker analysis workflows, it abstracts away the boilerplate of managing dataset splits and file paths. You define your protocol once in YAML, load it into the registry, and iterate over resources in Python—each resource is a dict-like object with keys populated from metadata files and custom loaders. The package depends on pandas, pyannote-core, and pyyaml, and requires Python 3.10 or later.
Use it for
- Organize speaker diarization datasets with train/dev/test splits and associated RTTM speaker annotations.
- Load speaker verification protocols with metadata from multiple file formats (RTTM, CTM, UEM) automatically selected by suffix.
- Define reproducible experimental protocols for audio segmentation tasks with lazy-loaded metadata and resource URIs.
- Build custom data loaders for proprietary audio or video metadata formats and register them with the protocol system.
- Iterate over multimedia resources with on-the-fly preprocessing and augmentation without modifying the underlying dataset files.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are working with speaker diarization, speaker verification, or other multimedia ML tasks and need a standardized way to define and iterate over dataset splits.
The low install friction and stable dependency set make it straightforward to adopt. However, verify the license terms first (they are not declared in the package metadata), and be aware that maintenance is aging—expect no rapid updates but likely sufficient stability for established workflows.
Install
pyannote-database on PyPI
Before you install
Low install friction with a pure-Python wheel and minimal dependencies (pandas, pyannote-core, pyyaml). Maintenance status is aging—last release was 250 days ago—so expect slower response to issues but likely stable for established use.
Requires Python 3.10 or later. YAML configuration file must be created and referenced; automatic loading checks ~/.pyannote/database.yml, ./database.yml, and PYANNOTE_DATABASE_CONFIG environment variable.
License in practice
License treatment is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.
Quickstart
pip install pyannote.database
from pyannote.database import registry
registry.load_database("/path/to/database.yml")
protocol = registry.get_protocol('MyDatabase.Protocol.MyProtocol')
for resource in protocol.train():
print(resource["uri"])
Verify before relying
- Whether the package is actively maintained or in long-term stable mode despite the 250-day release gap.
- What license actually governs this package (SPDX and raw license data are both missing).
- Whether custom data loaders and preprocessors cover common audio/video formats beyond RTTM, UEM, and CTM.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagespandaspyannote-corepyyaml |
| Maintenance | Aging 250 days since the last release |
| First released | |
| Downloads | 2,966,223 / month, #2,805 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: pyannote_database-6.1.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “multimedia database protocols”
- pyannote-databaseDefines and manages reproducible experimental protocols for…
- PySide6-AddonsPySide6-Addons extends PySide6 with additional Qt modules for 3D…
- pygletpyglet is a cross-platform windowing and multimedia library for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also pyannote-audio · pyannoteai-sdk · pyannote-metrics · pyannote-core · speechmatics-batch · lhotse · Resemblyzer · sherpa-onnx · mkdocs-video · pyannote-pipeline