$npx skillfedfor your agent

openai-whisper

Robust Speech Recognition via Large-Scale Weak Supervision

With conditionsPyPI Artificial IntelligenceReleased Jun 20254.2M downloads / moMITSource build

Decision gist · record as of 2026-08-14

sdist only — openai_whisper-20250625.tar.gz · builds from source
>>>import whispertop-level module
v20250625 · released 2025-06-26 · Python >=3.8 · 7 runtime deps: more-itertools, numba, numpy, tiktoken, torch, tqdm, triton

Yes, if you have the system dependencies and can tolerate high install friction. Whisper is actively maintained, permissively licensed, widely used (top 5000 packages), and has no known vulnerabilities. The heavy numerical dependencies (torch, numba, triton) are unavoidable for the task. Install only if you need multilingual speech recognition or translation; for English-only transcription, smaller or specialized models may be more practical.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires ffmpeg installed on your system (apt/pacman/brew/choco/scoop).
  • May also require Rust if tiktoken has no pre-built wheel for your platform.
  • High install friction: requires torch, numba, triton, and other heavy numerical dependencies.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects.

last release 2025-06-26 (414 days) · last repo commit 2026-07-28 · 107,269 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 4,241,268 downloads/mo, #2,355 on PyPI

Verify before relying

pip install openai-whisper

import whisper
model = whisper.load_model("turbo")
result = model.transcribe("audio.mp3")
print(result["text"])
  • Whether pre-built tiktoken wheels cover common platforms or if Rust compilation is frequently needed
  • Real-world transcription speed and accuracy on languages beyond English
  • Memory and compute requirements for models beyond the documented VRAM estimates
Same gist for agents: .md · .json

What it is and what it does

Whisper is OpenAI's general-purpose speech recognition model that transcribes, translates, and identifies languages in audio. It uses a Transformer sequence-to-sequence architecture trained on diverse multilingual audio data, allowing a single model to handle multiple speech-processing tasks that traditionally required separate pipeline stages. The model comes in six sizes (tiny, base, small, medium, large, turbo) with English-only and multilingual variants, offering speed-accuracy tradeoffs from ~1 GB to ~10 GB VRAM.

You can use it via command-line (e.g., `whisper audio.mp3 --model turbo`) or Python API. It processes audio in 30-second sliding windows and supports language specification and translation tasks. Installation requires torch, numba, triton, and other heavy numerical libraries, plus ffmpeg on your system. The package is actively maintained and has no known vulnerabilities.

Use it for

  • Transcribe English audio files quickly using the turbo model for real-time or batch processing
  • Translate non-English speech to English by specifying language and task parameters
  • Identify the spoken language in an audio file before further processing
  • Build a speech-to-text pipeline that handles multiple languages with a single model
  • Process audio in Python with fine-grained control via lower-level APIs like detect_language() and decode()

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you have the system dependencies and can tolerate high install friction.

Whisper is actively maintained, permissively licensed, widely used (top 5000 packages), and has no known vulnerabilities. The heavy numerical dependencies (torch, numba, triton) are unavoidable for the task. Install only if you need multilingual speech recognition or translation; for English-only transcription, smaller or specialized models may be more practical.

Install

openai-whisper on PyPI

Before you install

High install friction: requires torch, numba, triton, and other heavy numerical dependencies. Also requires ffmpeg as a system dependency and may need Rust installed if tiktoken lacks a pre-built wheel for your platform. Maintenance is active with recent commits.

Requires ffmpeg installed on your system (apt/pacman/brew/choco/scoop). May also require Rust if tiktoken has no pre-built wheel for your platform.

License in practice

MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects.

Quickstart

pip install openai-whisper

import whisper
model = whisper.load_model("turbo")
result = model.transcribe("audio.mp3")
print(result["text"])

Verify before relying

  • Whether pre-built tiktoken wheels cover common platforms or if Rust compilation is frequently needed
  • Real-world transcription speed and accuracy on languages beyond English
  • Memory and compute requirements for models beyond the documented VRAM estimates

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.8
Install frictionHigh. Source build required
Runtime dependencies
7 packages
more-itertoolsnumbanumpytiktokentorchtqdmtriton
MaintenanceActively maintained 414 days since the last release
Last repo commit
First released
Downloads4,241,268 / month, #2,355 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: openai_whisper-20250625.tar.gz

Tags

Capabilities
speech recognitionaudio transcriptionmultilingual speech-to-textspeech translationlanguage identification audiowhisper transcriptionaudio processing model
Topics
speech-recognitionmultilingualaudio-processing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “multilingual speech-to-text”

  • openai-whisperWhisper performs multilingual speech recognition, speech translation,…
  • qwen-asrQwen3-ASR provides speech recognition and language identification for…
  • whisper-timestampedAdds word-level timestamps and confidence scores to OpenAI's Whisper…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also faster-whisper · lhotse · mlx-whisper · SpeechRecognition · whisper-normalizer · vosk · whisper-timestamped · whisperx · pywhispercpp · omnivoice

Further reading