whisper-transcription
Turn podcasts, interviews, and videos into searchable text using OpenAI's Whisper model. Choose from multiple output formats—plain text, SRT subtitles, JSON, and more—with adjustable accuracy levels to match your needs. Handles batch processing and automatic language detection.
Whisper Transcription converts audio and video files to text using OpenAI's Whisper model for podcasts, interviews, and video subtitles.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-04-02
Whisper Transcription converts audio and video files to text using OpenAI's Whisper model for podcasts, interviews, and video subtitles. Turn podcasts, interviews, and videos into searchable text using OpenAI's Whisper model. Choose from multiple output formats—plain text, SRT subtitles, JSON, and more—with adjustable accuracy levels to match your needs. Handles batch processing and automatic language detection.
Use it when
- Yes, whisper-transcription generates video subtitles automatically in multiple formats including SRT.
- whisper-transcription supports batch processing of multiple audio files in formats like MP3 and MP4.
Verify before relying
Read SKILL.md below before installing (3 files). Open directory: indexed for reading, not audited.
Install
guia-matthieu/clawfu-skills/whisper-transcription · repository language: Python
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
How does whisper-transcription convert audio and video files to searchable, editable text transcripts?
whisper-transcription uses OpenAI's Whisper model to automatically convert audio and video files into searchable text transcripts. The tool processes your media, detects the language automatically, and outputs editable text with optional timestamps. You can adjust accuracy levels to balance speed and precision, making your recordings fully searchable and ready for editing or repurposing.
Can whisper-transcription generate video subtitles automatically?
Yes, whisper-transcription generates video subtitles automatically in multiple formats including SRT, which is compatible with most video platforms. The tool extracts audio from your video files, transcribes the content with timestamps, and outputs subtitle files ready for upload to YouTube, Vimeo, or other platforms. This makes adding captions to your videos quick and efficient.
What formats does whisper-transcription support for batch transcribing multiple audio files?
whisper-transcription supports batch processing of multiple audio files in formats like MP3 and MP4. You can process several files efficiently in one operation, with output available in plain text, SRT subtitles, JSON, and other formats. The batch feature is ideal for handling large volumes of podcasts, interviews, or recordings without processing them individually.
How can I convert a podcast to text for repurposing into a blog post?
whisper-transcription converts podcasts to text by extracting the audio, transcribing it with the Whisper model, and outputting editable text. You can then repurpose this transcript into blog posts, articles, or other written content. The tool's adjustable accuracy levels and multiple output formats make it easy to get clean, usable text from your podcast episodes.
Does whisper-transcription support transcribing and translating foreign language audio?
whisper-transcription includes automatic language detection and can transcribe foreign language recordings. While primary focus is on transcription, the tool's flexibility with multiple output formats and accuracy controls allows you to work with international audio content and create multilingual transcripts for global audiences.
What license does whisper-transcription use?
whisper-transcription is released under the MIT license, making it open-source and freely available for both personal and commercial use. The MIT license allows you to use, modify, and distribute the tool with minimal restrictions, provided you include the original license notice.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
Whisper Transcription
> Transcribe any audio or video to text using OpenAI's Whisper model - the same technology powering ChatGPT voice features.
When to Use This Skill
- Podcast repurposing - Convert episodes to blog posts, show notes, social snippets
- Video subtitles - Generate SRT/VTT files for YouTube, social media
- Interview extraction - Pull quotes and insights from recorded calls
- Content audit - Make audio/video libraries searchable
- Translation - Transcribe and translate foreign language content
What Claude Does vs What You Decide
| Claude Does | You Decide |
|---|---|
| Structures production workflow | Final creative direction |
| Suggests technical approaches | Equipment and tool choices |
| Creates templates and checklists | Quality |
(truncated - see the full file via the links below)
File tree — 3 files
skills/automation/whisper-transcription/SKILL.md
skills/automation/whisper-transcription/scripts/main.py
skills/automation/whisper-transcription/scripts/requirements.txt
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Convert audio and video files to searchable, editable text transcripts”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Video Processing harnesses FFmpeg to streamline common video workflows—compress for upload, extract audio, resize for social platforms, clip segments, merge files, and generate thumbnails. Ideal for content creators preparing videos across Instagram, TikTok, YouTube, and other platforms.
YouTube Downloader pulls videos, audio, and transcripts from YouTube using yt-dlp, supporting multiple formats and quality levels. Use it to archive webinars, extract audio for podcasts, grab transcripts for blogs, or gather competitor content for analysis. The skill handles individual videos, playlists, and metadata retrieval.
Content Repurposer breaks down long-form material—podcasts, articles, transcripts—into platform-specific short-form content like Twitter threads, LinkedIn posts, and Instagram carousels. It analyzes your source content, extracts key themes, and generates hooks, quotes, and formatted pieces ready to publish. Ideal for creators maximizing reach from a single piece of content.
Video Processor handles end-to-end video workflows: download from YouTube and thousands of other sites, convert between formats, pull audio tracks, and generate transcripts via Whisper. Built on yt-dlp, FFmpeg, and OpenAI's speech model.
Faster-whisper delivers rapid, offline speech-to-text transcription using CTranslate2, achieving 4-6x speed over OpenAI Whisper while maintaining identical accuracy. Generate subtitles in multiple formats (SRT, VTT, TTML, CSV), identify speakers, process batches with ETA, search transcripts, and detect chapters—all without API dependencies.
Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.
More skills video-summarizer (MIT) · whisper (MIT) · Whisper (Apache-2.0) · 9router-stt (MIT) · image-batch (MIT) · Speech To Text (Unlicense)