Funasr Transcribe
Funasr Transcribe processes audio and video files into structured Markdown transcripts with precise timestamps and speaker identification. It supports multiple formats (mp4, mov, mp3, wav, m4a, flac) and offers ONNX-accelerated recognition modes for faster processing. The skill automatically extracts video keyframes, generates AI-powered summaries, and handles both single-speaker and multi-speaker scenarios.
Funasr Transcribe converts audio and video files to timestamped text using local FunASR speech recognition.
AI-generated summary based on this skill's SKILL.md
Install
cat-xierluo/legal-skills/funasr-transcribe · repository language: Python
generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What audio formats does Funasr Transcribe support?
Funasr Transcribe processes multiple audio and video formats including mp4, mov, mp3, wav, m4a, and flac. The skill converts these files into structured Markdown transcripts with precise timestamps and speaker identification, making it easy to search and reference your audio content.
How do I transcribe audio to text with Funasr Transcribe?
Funasr Transcribe automatically converts your audio files into written text transcripts. Simply provide your audio or video file, and the skill performs automatic speech recognition, extracting the spoken content and organizing it with timestamps. For faster processing, you can use ONNX-accelerated recognition modes.
Can Funasr Transcribe handle multiple speakers?
Yes, Funasr Transcribe handles both single-speaker and multi-speaker scenarios. The skill identifies different speakers and includes speaker identification in your transcript, making it clear who said what throughout your audio content.
What additional features does Funasr Transcribe offer?
Beyond transcription, Funasr Transcribe automatically extracts video keyframes and generates AI-powered summaries of your content. These features help you quickly understand and navigate your audio and video files without listening to the entire recording.
Does Funasr Transcribe work with video files?
Funasr Transcribe processes both audio and video files. It transcribes the speech from video content and can extract keyframes, making it useful for converting video recordings, presentations, and multimedia content into searchable text transcripts with timestamps.
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Transcribe audio files to text using FunASR”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Convert local audio and video files into text transcripts using Whisper, then optionally generate meeting minutes, action items, speaker-separated output, and presentation-ready summaries. Supports Japanese meetings, Teams recordings, and other media formats entirely on your machine.
Release Workflow orchestrates the complete publication cycle for GitHub projects, from version synchronization through CI validation and artifact generation. It enforces cost-conscious practices to prevent misusing release pipelines as testing mechanisms, includes monorepo support for batch publishing multiple components, and provides recovery procedures for hotfixes.
MiMo V2.5 TTS transforms text into natural speech across three modes: preset voices for quick synthesis, voice design for custom tones via text description, and voice cloning from audio samples. Control emotion, dialect, and style through natural language, audio tags, or director mode for cinematic-quality output.
Feishu Voice TTS transforms text into speech using edge-tts and delivers it as Feishu audio messages, bypassing the platform's text fallback for direct file sends. The skill handles transcoding to Opus format, file upload, and message delivery through Feishu's open API.
Byted Voice to Text transcribes audio using Volcano Engine's BigModel ASR, offering fast synchronous processing for files under 2 hours and 100MB, or asynchronous recognition for longer content up to 5 hours. It handles Feishu voice messages, local audio files, and direct URLs with automatic format detection.
Byted-Text-to-Speech transforms written content into audio using Volcano Engine's speech synthesis service. Configure speaker voice, speech rate, pitch, volume, and output format—with optional Markdown filtering and LaTeX formula support for technical content.
More skills asr-transcribe-to-text (MIT) · Aj Patent Disclosure Cn (unlicensed) · Research Paper Writer (unlicensed)