{"enrichment":{"faq":[{"a":"Byted Voice to Text supports automatic format detection for common audio files including MP3, WAV, OGG, and other standard formats. The skill handles both local audio files and URLs, making it flexible for various transcription workflows.","q":"What audio formats does byted-voice-to-text support?"},{"a":"Yes, byted-voice-to-text uses Volcano Engine's BigModel ASR to transcribe voice to text recognition. It offers fast synchronous processing for files under 2 hours and 100MB, or asynchronous recognition for longer audio up to 5 hours.","q":"Can I transcribe voice to text recognition with byted-voice-to-text?"},{"a":"Byted Voice to Text can process Feishu voice messages automatically, converting them to text alongside support for local audio files and direct URLs. This integration streamlines voice message handling within Feishu workflows.","q":"Does byted-voice-to-text work with Feishu voice messages?"},{"a":"Byted Voice to Text provides asynchronous recognition for long-duration audio up to 5 hours, complementing its fast synchronous mode for shorter files. This dual approach ensures efficient processing regardless of audio length.","q":"How does byted-voice-to-text handle long audio transcription?"},{"a":"Yes, byted-voice-to-text converts speech from both URLs and local files to text quickly. It supports direct URL input with automatic format detection, enabling seamless transcription workflows.","q":"Can byted-voice-to-text transcribe audio from URLs?"},{"a":"Byted Voice to Text handles synchronous processing for files under 2 hours and 100MB. For longer content, use asynchronous recognition which supports audio up to 5 hours, providing flexibility for various transcription needs.","q":"What are the size and duration limits for byted-voice-to-text?"}],"shadow_tags":["audio-transcription","asr-engine","voice-recognition","multi-format-support","async-processing","feishu-integration","volcano-engine","real-time-conversion","multilingual-asr"],"summary_rewrite":"Byted Voice to Text transcribes audio using Volcano Engine's BigModel ASR, offering fast synchronous processing for files under 2 hours and 100MB, or asynchronous recognition for longer content up to 5 hours. It handles Feishu voice messages, local audio files, and direct URLs with automatic format detection."},"files":[{"bytes":7405,"path":"skills/byted-voice-to-text/SKILL.md","sha256":"393283cb55c5faf9bcebac51e66a80537c1b1b3c24f6457a889764e23a7d6eb1","url":"https://skillfed.io/files/bytedance/agentkit-samples/byted-voice-to-text/86ad4448/SKILL.md"}],"id":"bytedance/agentkit-samples/byted-voice-to-text","links":{"html":"https://skillfed.io/bytedance/agentkit-samples/byted-voice-to-text","md":"https://skillfed.io/bytedance/agentkit-samples/byted-voice-to-text.md","repo":"https://github.com/bytedance/agentkit-samples"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":84,"language":"Python","last_updated":"2026-07-27","license":"Apache-2.0","name":"byted-voice-to-text","publisher":"bytedance","stars":378},"relations":{"similar":[{"id":"bytedance/agentkit-samples/byted-text-to-speech"},{"id":"bytedance/agentkit-samples/byted-podcast-gen"},{"id":"bytedance/agentkit-samples/byted-mediakit-voiceover-editing"},{"id":"bytedance/agentkit-samples/byted-ark-seedance-guide"},{"id":"bytedance/agentkit-samples/volcengine-documentation"},{"id":"bytedance/agentkit-samples/byted-las-image-resample"},{"id":"bytedance/agentkit-samples/byted-vms-aicall"},{"id":"bytedance/agentkit-samples/byted-vms-secret-number"},{"id":"bytedance/agentkit-samples/byted-vod-process-tools"},{"id":"bytedance/agentkit-samples/byted-data-deepresearch-structured2markdown"}]},"slug":{"owner":"bytedance","repo":"agentkit-samples","skill":"byted-voice-to-text"},"version":"86ad4448"}
