Local Media Transcription
Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Local Media Transcription converts your audio and video files to text using local processing, with optional meeting notes and speaker labels.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-19
Local Media Transcription converts your audio and video files to text using local processing, with optional meeting notes and speaker labels. Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Use it when
- Local Media Transcription uses on-device Whisper processing to convert speech to text without requiring internet connectivity or external.
- Yes, Local Media Transcription transcribes MP3 and other media formats without uploading to the cloud.
Install
aktsmm/Agent-Skills/local-media-transcription · repository language: Python
generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
Can Local Media Transcription transcribe local media files?
Local Media Transcription converts audio and video files into text transcripts using Whisper, processing everything on your machine without cloud upload. The skill handles various media formats and supports Japanese meetings, Teams recordings, and other sources entirely on-device.
How does Local Media Transcription convert speech to text locally?
Local Media Transcription uses on-device Whisper processing to convert speech to text without requiring internet connectivity or external services. This on-device approach keeps your media private while delivering transcription results directly on your computer.
Can I transcribe mp3 without cloud upload using this skill?
Yes, Local Media Transcription transcribes MP3 and other media formats without uploading to the cloud. All processing happens locally on your machine, ensuring your audio files remain private and secure throughout the transcription process.
What can Local Media Transcription generate beyond basic transcripts?
Beyond transcription, Local Media Transcription can optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries from your transcribed content, all processed on your device.
Does Local Media Transcription work offline?
Local Media Transcription processes media files entirely offline on your machine, requiring no internet connectivity. This offline capability makes it ideal for private audio transcription and ensures your data never leaves your computer.
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Transcribe audio and video files locally without cloud upload”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Video Watch processes video URLs and local files to generate artifacts that GitHub Copilot can analyze, including captions, sampled frames, contact sheets, and a prompt packet. Use it when you need Copilot to inspect, summarize, or diagnose video content like demos, recordings, or tutorials. The skill handles multiple detail modes to balance speed and visual coverage.
Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.
Whisper is OpenAI's multilingual speech recognition model for converting audio and video into text across 99 languages. It handles noisy recordings, supports translation to English, and offers multiple model sizes from lightweight to high-accuracy variants. Use it for podcasts, meeting transcription, video subtitles, and multilingual audio processing.
Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.
Media Transcription runs a durable, event-driven pipeline to transcribe meeting media stored on the NAS, using MLX Whisper for speech recognition and pyannote for speaker identification. Monitor progress in real time, cancel runs, or resume from partial completions—all orchestrated through Inngest with detached local inference processes that prevent timeout failures. Execution is limited to the Flagg host worker with direct access to the media mount.
Convert audio and video files into transcribed text with automatic speaker identification and timing markers. Choose between local processing on Apple Silicon Macs or remote API endpoints, with optional speaker diarization and plain-text fallback.
More skills Funasr Transcribe (unlicensed) · video-processor (MIT) · interview-transcription (MIT) · Context To Video (NOASSERTION)