$npx skillfedfor your agent

whisper-transcription

Turn podcasts, interviews, and videos into searchable text using OpenAI's Whisper model. Choose from multiple output formats—plain text, SRT subtitles, JSON, and more—with adjustable accuracy levels to match your needs. Handles batch processing and automatic language detection.

Whisper Transcription converts audio and video files to text using OpenAI's Whisper model for podcasts, interviews, and video subtitles.

AI-generated summary based on this skill's SKILL.md

140 28 MITupdated by guia-matthieu

Decision gist · record as of 2026-04-02

Whisper Transcription converts audio and video files to text using OpenAI's Whisper model for podcasts, interviews, and video subtitles. Turn podcasts, interviews, and videos into searchable text using OpenAI's Whisper model. Choose from multiple output formats—plain text, SRT subtitles, JSON, and more—with adjustable accuracy levels to match your needs. Handles batch processing and automatic language detection.

manual: git clone https://github.com/guia-matthieu/clawfu-skills → cp -r clawfu-skills/skills/automation/whisper-transcription ~/.claude/skills/whisper-transcription
skills/automation/whisper-transcription/SKILL.md · version 2fe4f038

Use it when

  • Yes, whisper-transcription generates video subtitles automatically in multiple formats including SRT.
  • whisper-transcription supports batch processing of multiple audio files in formats like MP3 and MP4.

Verify before relying

Read SKILL.md below before installing (3 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

guia-matthieu/clawfu-skills/whisper-transcription · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

How does whisper-transcription convert audio and video files to searchable, editable text transcripts?

whisper-transcription uses OpenAI's Whisper model to automatically convert audio and video files into searchable text transcripts. The tool processes your media, detects the language automatically, and outputs editable text with optional timestamps. You can adjust accuracy levels to balance speed and precision, making your recordings fully searchable and ready for editing or repurposing.

Can whisper-transcription generate video subtitles automatically?

Yes, whisper-transcription generates video subtitles automatically in multiple formats including SRT, which is compatible with most video platforms. The tool extracts audio from your video files, transcribes the content with timestamps, and outputs subtitle files ready for upload to YouTube, Vimeo, or other platforms. This makes adding captions to your videos quick and efficient.

What formats does whisper-transcription support for batch transcribing multiple audio files?

whisper-transcription supports batch processing of multiple audio files in formats like MP3 and MP4. You can process several files efficiently in one operation, with output available in plain text, SRT subtitles, JSON, and other formats. The batch feature is ideal for handling large volumes of podcasts, interviews, or recordings without processing them individually.

How can I convert a podcast to text for repurposing into a blog post?

whisper-transcription converts podcasts to text by extracting the audio, transcribing it with the Whisper model, and outputting editable text. You can then repurpose this transcript into blog posts, articles, or other written content. The tool's adjustable accuracy levels and multiple output formats make it easy to get clean, usable text from your podcast episodes.

Does whisper-transcription support transcribing and translating foreign language audio?

whisper-transcription includes automatic language detection and can transcribe foreign language recordings. While primary focus is on transcription, the tool's flexibility with multiple output formats and accuracy controls allows you to work with international audio content and create multilingual transcripts for global audiences.

What license does whisper-transcription use?

whisper-transcription is released under the MIT license, making it open-source and freely available for both personal and commercial use. The MIT license allows you to use, modify, and distribute the tool with minimal restrictions, provided you include the original license notice.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

Whisper Transcription

> Transcribe any audio or video to text using OpenAI's Whisper model - the same technology powering ChatGPT voice features.

When to Use This Skill

  • Podcast repurposing - Convert episodes to blog posts, show notes, social snippets
  • Video subtitles - Generate SRT/VTT files for YouTube, social media
  • Interview extraction - Pull quotes and insights from recorded calls
  • Content audit - Make audio/video libraries searchable
  • Translation - Transcribe and translate foreign language content

What Claude Does vs What You Decide

Claude Does You Decide
Structures production workflow Final creative direction
Suggests technical approaches Equipment and tool choices
Creates templates and checklists Quality

(truncated - see the full file via the links below)

File tree — 3 files
skills/automation/whisper-transcription/SKILL.md
skills/automation/whisper-transcription/scripts/main.py
skills/automation/whisper-transcription/scripts/requirements.txt

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert audio and video files to searchable, editable text transcripts”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

video-processing
by guia-matthieu · guia-matthieu/clawfu-skills

Video Processing harnesses FFmpeg to streamline common video workflows—compress for upload, extract audio, resize for social platforms, clip segments, merge files, and generate thumbnails. Ideal for content creators preparing videos across Instagram, TikTok, YouTube, and other platforms.

MITupdated Apr 2026
★ 140repo stars
youtube-downloader
by guia-matthieu · guia-matthieu/clawfu-skills

YouTube Downloader pulls videos, audio, and transcripts from YouTube using yt-dlp, supporting multiple formats and quality levels. Use it to archive webinars, extract audio for podcasts, grab transcripts for blogs, or gather competitor content for analysis. The skill handles individual videos, playlists, and metadata retrieval.

MITupdated Apr 2026
★ 140repo stars
content-repurposer
by guia-matthieu · guia-matthieu/clawfu-skills

Content Repurposer breaks down long-form material—podcasts, articles, transcripts—into platform-specific short-form content like Twitter threads, LinkedIn posts, and Instagram carousels. It analyzes your source content, extracts key themes, and generates hooks, quotes, and formatted pieces ready to publish. Ideal for creators maximizing reach from a single piece of content.

MITupdated Apr 2026
★ 140repo stars
video-processor
by iamzhihuix · iamzhihuix/happy-claude-skills

Video Processor handles end-to-end video workflows: download from YouTube and thousands of other sites, convert between formats, pull audio tracks, and generate transcripts via Whisper. Built on yt-dlp, FFmpeg, and OpenAI's speech model.

MITupdated Apr 2026
★ 305repo stars
faster-whisper
by ThePlasmak · ThePlasmak/faster-whisper

Faster-whisper delivers rapid, offline speech-to-text transcription using CTranslate2, achieving 4-6x speed over OpenAI Whisper while maintaining identical accuracy. Generate subtitles in multiple formats (SRT, VTT, TTML, CSV), identify speakers, process batches with ETA, search transcripts, and detect chapters—all without API dependencies.

MITupdated Feb 2026
★ 9repo stars
Transcription Automation
by claude-office-skills · claude-office-skills/skills

Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.

MITupdated Jan 2026
★ 338repo stars

More skills video-summarizer (MIT) · whisper (MIT) · Whisper (Apache-2.0) · 9router-stt (MIT) · image-batch (MIT) · Speech To Text (Unlicense)

Tags
speech-to-textmedia-conversioncontent-repurposingbatch-processingsubtitle-generationaudio-indexingmultilingual-supporttimestamp-extractionworkflow-automation