Whisper
Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.
Whisper transcribes speech and audio files to text across 99 languages with six model sizes.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-26
Whisper transcribes speech and audio files to text across 99 languages with six model sizes. Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.
Use it when
- Yes.
- Whisper processes podcasts, meetings, and video audio into searchable transcripts with optional timestamped segments.
Verify before relying
Read SKILL.md below before installing (2 files). Open directory: indexed for reading, not audited.
Install
graniet/kheish/whisper · repository language: Rust
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What is Whisper speech to text and how does it work?
Whisper is OpenAI's multilingual automatic speech recognition model trained on 680,000 hours of audio data. It transcribes speech to text across 99 languages, identifies the spoken language automatically, and can translate audio content to English. Whisper offers six configurable model sizes ranging from 39M to 1550M parameters, letting you balance accuracy against computational cost.
Can Whisper convert audio to text in multiple languages?
Yes. Whisper supports transcription and translation across 99 languages. You can transcribe audio in any of those languages to text in the original language, or translate the audio content to English. The model automatically detects which language is being spoken, making it ideal for multilingual podcasts, meetings, and international content.
How can I use Whisper for podcast transcription and meeting notes?
Whisper processes podcasts, meetings, and video audio into searchable transcripts with optional timestamped segments. Its robust handling of noisy audio, background noise, and varied speaking styles makes it well-suited for real-world recordings. You can run Whisper locally on your hardware or integrate it into transcription workflows without relying on cloud APIs.
Is Whisper open source and can I run it locally?
Whisper is released under the Apache-2.0 license and available as open-source software. You can run it locally on your own hardware without cloud API costs or dependencies on external services. The model supports GPU acceleration for faster processing, and its six size options let you choose the right trade-off between speed and accuracy for your setup.
What audio formats and file types does Whisper support?
Whisper handles common audio formats including MP3, WAV, M4A, FLAC, and others. It can process audio from files, video content, or streaming sources. The model is designed to handle real-world audio quality, including background noise and poor recording conditions, making it practical for transcribing podcasts, meetings, and archival material.
How do Whisper's different model sizes affect performance?
Whisper offers six model sizes from 39M to 1550M parameters. Smaller models run faster and use less memory, making them suitable for real-time or resource-constrained environments. Larger models deliver higher accuracy, especially on challenging audio. Choose based on your hardware, latency requirements, and accuracy needs—GPU acceleration can speed up processing across all sizes.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
Kheish Compatibility
This skill is repo-local and stays inactive until explicitly activated.
When the original instructions refer to legacy tool names, use these Kheish mappings:
terminal=>bashweb_extract=>web_fetch, plusweb_searchwhen discovery is neededsearch_files=>grep_searchandglob_searchbrowser_*tools require a browser-capable surfaced tool or MCP; if none is available, use the closest available surface and say so explicitly
When the instructions mention local helper
(truncated - see the full file via the links below)
File tree — 2 files
skills/mlops/models/whisper/SKILL.md
skills/mlops/models/whisper/references/languages.md
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Transcribe speech or audio files to text using multilingual ASR”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Whisper is OpenAI's multilingual speech recognition model for converting audio and video into text across 99 languages. It handles noisy recordings, supports translation to English, and offers multiple model sizes from lightweight to high-accuracy variants. Use it for podcasts, meeting transcription, video subtitles, and multilingual audio processing.
Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Faster-whisper delivers rapid, offline speech-to-text transcription using CTranslate2, achieving 4-6x speed over OpenAI Whisper while maintaining identical accuracy. Generate subtitles in multiple formats (SRT, VTT, TTML, CSV), identify speakers, process batches with ETA, search transcripts, and detect chapters—all without API dependencies.
Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.
This skill implements speech-to-text using Faster Whisper for converting audio input into written transcriptions. It prioritizes local processing, immediate deletion of audio data, and secure handling of voice information while supporting real-time streaming, multiple languages, and hardware-optimized model selection.
Turn podcasts, interviews, and videos into searchable text using OpenAI's Whisper model. Choose from multiple output formats—plain text, SRT subtitles, JSON, and more—with adjustable accuracy levels to match your needs. Handles batch processing and automatic language detection.
More skills 9router-stt (MIT) · happy-audio-gen (MIT)