$npx skillfedfor your agent

Local Media Transcription

Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.

Local Media Transcription converts your audio and video files to text using local processing, with optional meeting notes and speaker labels.

AI-generated summary based on this skill's SKILL.md

22 4 unlicensed, metadata onlyupdated by aktsmm

Decision gist · record as of 2026-07-19

Local Media Transcription converts your audio and video files to text using local processing, with optional meeting notes and speaker labels. Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.

manual: git clone https://github.com/aktsmm/Agent-Skills → cp -r Agent-Skills ~/.claude/skills/local-media-transcription

Use it when

  • Local Media Transcription uses on-device Whisper processing to convert speech to text without requiring internet connectivity or external.
  • Yes, Local Media Transcription transcribes MP3 and other media formats without uploading to the cloud.
Same gist for agents: .md · .json

Install

aktsmm/Agent-Skills/local-media-transcription · repository language: Python

generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

Can Local Media Transcription transcribe local media files?

Local Media Transcription converts audio and video files into text transcripts using Whisper, processing everything on your machine without cloud upload. The skill handles various media formats and supports Japanese meetings, Teams recordings, and other sources entirely on-device.

How does Local Media Transcription convert speech to text locally?

Local Media Transcription uses on-device Whisper processing to convert speech to text without requiring internet connectivity or external services. This on-device approach keeps your media private while delivering transcription results directly on your computer.

Can I transcribe mp3 without cloud upload using this skill?

Yes, Local Media Transcription transcribes MP3 and other media formats without uploading to the cloud. All processing happens locally on your machine, ensuring your audio files remain private and secure throughout the transcription process.

What can Local Media Transcription generate beyond basic transcripts?

Beyond transcription, Local Media Transcription can optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries from your transcribed content, all processed on your device.

Does Local Media Transcription work offline?

Local Media Transcription processes media files entirely offline on your machine, requiring no internet connectivity. This offline capability makes it ideal for private audio transcription and ensures your data never leaves your computer.

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Transcribe audio and video files locally without cloud upload”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

Video Watch
by aktsmm · aktsmm/Agent-Skills

Video Watch processes video URLs and local files to generate artifacts that GitHub Copilot can analyze, including captions, sampled frames, contact sheets, and a prompt packet. Use it when you need Copilot to inspect, summarize, or diagnose video content like demos, recordings, or tutorials. The skill handles multiple detail modes to balance speed and visual coverage.

no license declared → metadata onlyupdated Jul 2026
★ 22repo stars
Transcription Automation
by claude-office-skills · claude-office-skills/skills

Transcription Automation handles speech-to-text conversion for audio files, video recordings, and live streams, automatically identifying speakers and generating formatted transcripts. The skill produces searchable archives, meeting notes with action items, and subtitles in SRT or VTT formats across multiple languages. It integrates with platforms like Zoom, YouTube, and podcasting workflows to streamline content processing end-to-end.

MITupdated Jan 2026
★ 338repo stars
whisper
by NousResearch · NousResearch/hermes-agent

Whisper is OpenAI's multilingual speech recognition model for converting audio and video into text across 99 languages. It handles noisy recordings, supports translation to English, and offers multiple model sizes from lightweight to high-accuracy variants. Use it for podcasts, meeting transcription, video subtitles, and multilingual audio processing.

MITupdated Jul 2026
★ 221,503repo stars
Whisper
by graniet · graniet/kheish

Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.

Apache-2.0updated Jul 2026
★ 264repo stars
Media Transcription
by joelhooks · joelhooks/joelclaw

Media Transcription runs a durable, event-driven pipeline to transcribe meeting media stored on the NAS, using MLX Whisper for speech recognition and pyannote for speaker identification. Monitor progress in real time, cancel runs, or resume from partial completions—all orchestrated through Inngest with detached local inference processes that prevent timeout failures. Execution is limited to the Flagg host worker with direct access to the media mount.

no license declared → metadata onlyupdated Jul 2026
★ 61repo stars
asr-transcribe-to-text
by daymade · daymade/claude-code-skills

Convert audio and video files into transcribed text with automatic speaker identification and timing markers. Choose between local processing on Apple Silicon Macs or remote API endpoints, with optional speaker diarization and plain-text fallback.

MITupdated Jul 2026
★ 1,299repo stars

More skills Funasr Transcribe (unlicensed) · video-processor (MIT) · interview-transcription (MIT) · Context To Video (NOASSERTION)

Tags
offline-processingprivacy-focusedon-device-aiaudio-conversionspeech-recognitionlocal-storageno-cloud-uploadmedia-processing