Transcription Automation
Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.
Transcription Automation converts audio and video files into searchable text transcripts with speaker identification and subtitle generation.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-01-31
Transcription Automation converts audio and video files into searchable text transcripts with speaker identification and subtitle generation. Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.
Use it when
- Transcription Automation automates meeting recording transcription by identifying and labeling different speakers throughout the recording.
- Yes, Transcription Automation automatically generates subtitles and captions for video content in standard formats like SRT and VTT.
Verify before relying
Read SKILL.md below before installing (1 file). Open directory: indexed for reading, not audited.
Install
claude-office-skills/skills/transcription-automation · repository language: TypeScript
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What can Transcription Automation do with audio files and video recordings?
Transcription Automation converts speech-to-text from audio files, video recordings, and live streams into searchable, formatted transcripts. It automatically identifies speakers, adds timestamps, and generates subtitles in SRT or VTT formats. The skill produces organized meeting notes with action items and supports batch processing of multiple files simultaneously.
How does Transcription Automation handle meeting recordings with speaker labels?
Transcription Automation automates meeting recording transcription by identifying and labeling different speakers throughout the recording. This speaker diarization feature creates transcripts that clearly show who said what, making it easy to track contributions and extract meeting notes. The skill integrates with platforms like Zoom to streamline the entire workflow.
Can Transcription Automation auto generate subtitles for videos?
Yes, Transcription Automation automatically generates subtitles and captions for video content in standard formats like SRT and VTT. It processes YouTube videos, podcasts, and other video files to produce searchable text with precise timestamps. This makes video content accessible and enables full-text search across your video library.
Does Transcription Automation support batch transcribe multiple audio files?
Transcription Automation supports batch processing to transcribe multiple audio files efficiently. You can process large volumes of content at once, making it ideal for podcast archives, meeting backlogs, and bulk content libraries. Each file receives speaker identification, timestamps, and formatted transcripts for easy organization and search.
What languages does Transcription Automation support in transcription workflows?
Transcription Automation provides multilingual transcription support across multiple languages, enabling global workflows. Whether you're processing international meetings, multilingual podcasts, or diverse video content, the skill handles language detection and transcription to create searchable archives in your required languages.
How does Transcription Automation extract and organize meeting notes?
Transcription Automation extracts and organizes meeting notes from recordings by transcribing speech, identifying speakers, and structuring content with timestamps and action items. The skill integrates with meeting platforms to automate note-taking, making it simple to archive discussions, track decisions, and follow up on action items without manual transcription.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
Transcription Automation
Comprehensive skill for automating audio/video transcription and content processing.
Core Workflows
1. Transcription Pipeline
``` TRANSCRIPTION FLOW: ┌─────────────────┐ │ Audio/Video │ │ Input │ └────────┬────────┘ ▼ ┌─────────────────┐ │
(truncated - see the full file via the links below)
File tree — 1 file
transcription-automation/SKILL.md
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Transcribe audio and video files to searchable text”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Faster-whisper delivers rapid, offline speech-to-text transcription using CTranslate2, achieving 4-6x speed over OpenAI Whisper while maintaining identical accuracy. Generate subtitles in multiple formats (SRT, VTT, TTML, CSV), identify speakers, process batches with ETA, search transcripts, and detect chapters—all without API dependencies.
Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Whisper is OpenAI's multilingual speech recognition model for converting audio and video into text across 99 languages. It handles noisy recordings, supports translation to English, and offers multiple model sizes from lightweight to high-accuracy variants. Use it for podcasts, meeting transcription, video subtitles, and multilingual audio processing.
Convert interview recordings into searchable transcripts with word-level timestamps and speaker identification. Extract and verify quotes for publication, track sources, and organize multi-interview projects with built-in templates for manual transcription and quote management.
Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.
Convert audio and video files into transcribed text with automatic speaker identification and timing markers. Choose between local processing on Apple Silicon Macs or remote API endpoints, with optional speaker diarization and plain-text fallback.
More skills whisper-transcription (MIT) · discord-bot (MIT) · transcript-polisher (MIT) · doubao-tts (MIT) · baoyu-youtube-transcript (MIT) · Trello Automation (MIT) · ClickUp Automation (MIT)