$npx skillfedfor your agent

Transcription Automation

Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.

Transcription Automation converts audio and video files into searchable text transcripts with speaker identification and subtitle generation.

AI-generated summary based on this skill's SKILL.md

338 71 MITupdated by claude-office-skills

Decision gist · record as of 2026-01-31

Transcription Automation converts audio and video files into searchable text transcripts with speaker identification and subtitle generation. Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.

manual: git clone https://github.com/claude-office-skills/skills → cp -r skills/transcription-automation ~/.claude/skills/transcription-automation
transcription-automation/SKILL.md · version 55ba688b

Use it when

  • Transcription Automation automates meeting recording transcription by identifying and labeling different speakers throughout the recording.
  • Yes, Transcription Automation automatically generates subtitles and captions for video content in standard formats like SRT and VTT.

Verify before relying

Read SKILL.md below before installing (1 file). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

claude-office-skills/skills/transcription-automation · repository language: TypeScript

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What can Transcription Automation do with audio files and video recordings?

Transcription Automation converts speech-to-text from audio files, video recordings, and live streams into searchable, formatted transcripts. It automatically identifies speakers, adds timestamps, and generates subtitles in SRT or VTT formats. The skill produces organized meeting notes with action items and supports batch processing of multiple files simultaneously.

How does Transcription Automation handle meeting recordings with speaker labels?

Transcription Automation automates meeting recording transcription by identifying and labeling different speakers throughout the recording. This speaker diarization feature creates transcripts that clearly show who said what, making it easy to track contributions and extract meeting notes. The skill integrates with platforms like Zoom to streamline the entire workflow.

Can Transcription Automation auto generate subtitles for videos?

Yes, Transcription Automation automatically generates subtitles and captions for video content in standard formats like SRT and VTT. It processes YouTube videos, podcasts, and other video files to produce searchable text with precise timestamps. This makes video content accessible and enables full-text search across your video library.

Does Transcription Automation support batch transcribe multiple audio files?

Transcription Automation supports batch processing to transcribe multiple audio files efficiently. You can process large volumes of content at once, making it ideal for podcast archives, meeting backlogs, and bulk content libraries. Each file receives speaker identification, timestamps, and formatted transcripts for easy organization and search.

What languages does Transcription Automation support in transcription workflows?

Transcription Automation provides multilingual transcription support across multiple languages, enabling global workflows. Whether you're processing international meetings, multilingual podcasts, or diverse video content, the skill handles language detection and transcription to create searchable archives in your required languages.

How does Transcription Automation extract and organize meeting notes?

Transcription Automation extracts and organizes meeting notes from recordings by transcribing speech, identifying speakers, and structuring content with timestamps and action items. The skill integrates with meeting platforms to automate note-taking, making it simple to archive discussions, track decisions, and follow up on action items without manual transcription.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

Transcription Automation

Comprehensive skill for automating audio/video transcription and content processing.

Core Workflows

1. Transcription Pipeline

``` TRANSCRIPTION FLOW: ┌─────────────────┐ │ Audio/Video │ │ Input │ └────────┬────────┘ ▼ ┌─────────────────┐ │

(truncated - see the full file via the links below)

File tree — 1 file
transcription-automation/SKILL.md

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Transcribe audio and video files to searchable text”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

faster-whisper
by ThePlasmak · ThePlasmak/faster-whisper

Faster-whisper delivers rapid, offline speech-to-text transcription using CTranslate2, achieving 4-6x speed over OpenAI Whisper while maintaining identical accuracy. Generate subtitles in multiple formats (SRT, VTT, TTML, CSV), identify speakers, process batches with ETA, search transcripts, and detect chapters—all without API dependencies.

MITupdated Feb 2026
★ 9repo stars
Local Media Transcription
by aktsmm · aktsmm/Agent-Skills

Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.

no license declared → metadata onlyupdated Jul 2026
★ 22repo stars
whisper
by NousResearch · NousResearch/hermes-agent

Whisper is OpenAI's multilingual speech recognition model for converting audio and video into text across 99 languages. It handles noisy recordings, supports translation to English, and offers multiple model sizes from lightweight to high-accuracy variants. Use it for podcasts, meeting transcription, video subtitles, and multilingual audio processing.

MITupdated Jul 2026
★ 221,503repo stars
interview-transcription
by jamditis · jamditis/claude-skills-journalism

Convert interview recordings into searchable transcripts with word-level timestamps and speaker identification. Extract and verify quotes for publication, track sources, and organize multi-interview projects with built-in templates for manual transcription and quote management.

MITupdated Jul 2026
★ 342repo stars
Whisper
by graniet · graniet/kheish

Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.

Apache-2.0updated Jul 2026
★ 264repo stars
asr-transcribe-to-text
by daymade · daymade/claude-code-skills

Convert audio and video files into transcribed text with automatic speaker identification and timing markers. Choose between local processing on Apple Silicon Macs or remote API endpoints, with optional speaker diarization and plain-text fallback.

MITupdated Jul 2026
★ 1,299repo stars

More skills whisper-transcription (MIT) · discord-bot (MIT) · transcript-polisher (MIT) · doubao-tts (MIT) · baoyu-youtube-transcript (MIT) · Trello Automation (MIT) · ClickUp Automation (MIT)

Tags
speech-recognitionmedia-processingmeeting-automationcontent-indexingspeaker-identificationsubtitle-generationmultilingual-supportworkflow-integrationquality-metrics