whisper-timestamped
Multi-lingual Automatic Speech Recognition (ASR) based on Whisper models, with accurate word timestamps, access to language detection confidence, several options for Voice Activity Detection (VAD), and more.
Decision gist · record as of 2026-08-14
Yes, if you need word-level timestamps for Whisper transcriptions and accept GPLv3 copyleft constraints. The package has low install friction, active maintenance, and no known vulnerabilities. Suitable for open-source projects, research, and accessibility tools; less suitable for proprietary closed-source products due to license. Aging status (339 days since release) is acceptable given recent repository activity.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires ffmpeg installed on system; Python >=3.7 (3.9+ recommended); openai-whisper must be installed as a runtime dependency.
- Low friction installation via pip; depends on Cython, dtw-python, and openai-whisper.
- Package is aging (339 days since last release) but repository remains active with 2837 stars and recent commits; maintenance signal is moderate.
License · maintenance · safety
GPLv3 (copyleft) — GPLv3 copyleft license means any derivative work or distribution must also be open-source under GPLv3; suitable for open-source projects but requires careful consideration in proprietary or closed-source contexts.
last release 2025-09-09 (339 days) · last repo commit 2025-09-09 · 2,837 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 98,887 downloads/mo, #13,052 on PyPI
Alternatives
Verify before relying
pip install whisper-timestamped
import whisper_timestamped as whisper
result = whisper.transcribe(model, "audio.mp3")- Whether word-timestamp accuracy remains within stated bounds across all supported languages and audio conditions.
- Performance impact of DTW alignment on inference time for typical audio lengths.
- Compatibility guarantees with future versions of openai-whisper beyond the stated 'any version' claim.
What it is and what it does
whisper-timestamped extends OpenAI's Whisper speech recognition model to predict word-level timestamps and confidence scores for each word in a transcription. Rather than Whisper's native segment-level timestamps (typically 1-second accuracy), this package uses Dynamic Time Warping applied to the model's cross-attention weights to align words precisely with their spoken timing. It processes long audio files with minimal additional memory overhead and works across Whisper's supported languages.
The package integrates optional Voice Activity Detection (VAD) to reduce hallucinations on silence, provides language detection confidence when the language is unspecified, and aims to do word alignment without extra inference steps when possible. It is designed as a drop-in extension to openai-whisper, maintaining API compatibility while adding these timing and confidence features for applications requiring precise word-level synchronization.
Use it for
- Generate subtitle files with accurate word-level timing for video synchronization and accessibility.
- Extract precise timestamps for each word to enable interactive transcript playback or speaker diarization.
- Detect and filter out hallucinated speech on silence using VAD before transcription to improve accuracy.
- Build searchable transcripts where users can click a word to jump to its exact position in audio.
- Analyze speech patterns by correlating word timing with confidence scores to identify disfluencies or hesitations.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need word-level timestamps for Whisper transcriptions and accept GPLv3 copyleft constraints.
The package has low install friction, active maintenance, and no known vulnerabilities. Suitable for open-source projects, research, and accessibility tools; less suitable for proprietary closed-source products due to license. Aging status (339 days since release) is acceptable given recent repository activity.
Install
whisper-timestamped on PyPI
Before you install
Low friction installation via pip; depends on Cython, dtw-python, and openai-whisper. Package is aging (339 days since last release) but repository remains active with 2837 stars and recent commits; maintenance signal is moderate.
Requires ffmpeg installed on system; Python >=3.7 (3.9+ recommended); openai-whisper must be installed as a runtime dependency.
License in practice
GPLv3 copyleft license means any derivative work or distribution must also be open-source under GPLv3; suitable for open-source projects but requires careful consideration in proprietary or closed-source contexts.
Quickstart
pip install whisper-timestamped
import whisper_timestamped as whisper
result = whisper.transcribe(model, "audio.mp3")
Verify before relying
- Whether word-timestamp accuracy remains within stated bounds across all supported languages and audio conditions.
- Performance impact of DTW alignment on inference time for typical audio lengths.
- Compatibility guarantees with future versions of openai-whisper beyond the stated 'any version' claim.
Package facts
| License | GPLv3 copyleft |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesCythondtw-pythonopenai-whisper |
| Maintenance | Aging 339 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 98,887 / month, #13,052 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: whisper_timestamped-1.15.9-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “word-level timestamps speech recognition”
- whisper-timestampedAdds word-level timestamps and confidence scores to OpenAI's Whisper…
- whisperxWhisperX performs fast automatic speech recognition with word-level…
- qwen-asrQwen3-ASR provides speech recognition and language identification for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also whisperx · openai-whisper · whisper · faster-whisper · realtimestt · funasr · whisper-normalizer · pvporcupine · mlx-whisper · SpeechRecognition