Media Transcription
Media Transcription runs a durable, event-driven pipeline to transcribe meeting media stored on the NAS, using MLX Whisper for speech recognition and pyannote for speaker identification. Monitor progress in real time, cancel runs, or resume from partial completions—all orchestrated through Inngest with detached local inference processes that prevent timeout failures. Execution is limited to the Flagg host worker with direct access to the media mount.
Media Transcription converts meeting audio and video files into searchable text transcripts using durable, event-driven processing.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-27
Media Transcription converts meeting audio and video files into searchable text transcripts using durable, event-driven processing. Media Transcription runs a durable, event-driven pipeline to transcribe meeting media stored on the NAS, using MLX Whisper for speech recognition and pyannote for speaker identification. Monitor progress in real time, cancel runs, or resume from partial completions—all orchestrated through Inngest with detached local inference processes that prevent timeout failures. Execution is limited to the Flagg host worker with direct access to the media mount.
Use it when
- Yes.
- Media Transcription lets you monitor progress in real time through its orchestration interface.
Similar skills
Install
joelhooks/joelclaw/media-transcription · repository language: TypeScript
generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What does Media Transcription do?
Media Transcription runs an event-driven pipeline to convert audio and video media files into searchable text transcripts. It uses MLX Whisper for speech recognition and pyannote for speaker identification, automatically extracting and processing spoken content from your media files to generate accessible text versions.
Can I transcribe video to text with Media Transcription?
Yes. Media Transcription transcribes video to text by processing media stored on the NAS. It handles video files like MP4s alongside audio formats, extracting speech and converting it into searchable transcripts with speaker identification included.
How do I monitor and control transcription runs?
Media Transcription lets you monitor progress in real time through its orchestration interface. You can cancel active runs or resume from partial completions, giving you full control over the transcription pipeline without losing work if a process is interrupted.
What audio and video formats does Media Transcription support?
Media Transcription processes media files stored on the NAS, including MP3 audio files, MP4 video files, and other common formats. The pipeline automatically handles format conversion as needed during the transcription workflow.
How does Media Transcription prevent timeout failures?
Media Transcription uses Inngest orchestration with detached local inference processes that run independently of the main pipeline. This architecture prevents timeout failures by allowing speech recognition and speaker identification to complete without being constrained by request timeouts.
Where does Media Transcription run and what access does it need?
Media Transcription execution is limited to the Flagg host worker, which has direct access to the media mount on the NAS. This ensures reliable, local processing of your media files with proper file system permissions and consistent performance.
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Convert audio or video media files into searchable text transcripts”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Wzrrd Video automates video publishing to wzrrd.sh watch pages with built-in player, transcript display, and automatic subtitle generation in six languages. The skill integrates with flagg's joelclaw door for direct publishing and supports the wzrrd CLI for uploads from any machine. Monitor encode progress, manage access, and embed videos in Brain notes with status tracking and revocation controls.
Convert audio and video files into transcribed text with automatic speaker identification and timing markers. Choose between local processing on Apple Silicon Macs or remote API endpoints, with optional speaker diarization and plain-text fallback.
Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Convert interview recordings into searchable transcripts with word-level timestamps and speaker identification. Extract and verify quotes for publication, track sources, and organize multi-interview projects with built-in templates for manual transcription and quote management.
Video Processor handles end-to-end video workflows: download from YouTube and thousands of other sites, convert between formats, pull audio tracks, and generate transcripts via Whisper. Built on yt-dlp, FFmpeg, and OpenAI's speech model.
YouTube Transcript extracts text content from videos by downloading available captions or auto-generated subtitles via yt-dlp. When subtitles aren't available, it can transcribe audio using Whisper as a fallback option.
More skills wjs-reframing-video (MIT)