$npx skillfedfor your agent

byted-text-to-speech

Byted-Text-to-Speech transforms written content into audio using Volcano Engine's speech synthesis service. Configure speaker voice, speech rate, pitch, volume, and output format—with optional Markdown filtering and LaTeX formula support for technical content.

Byted-Text-to-Speech converts text to speech using Volcano Engine's synthesis API with adjustable voice, speed, and tone.

AI-generated summary based on this skill's SKILL.md

378 84 Apache-2.0updated by bytedance

Decision gist · record as of 2026-07-27

Byted-Text-to-Speech converts text to speech using Volcano Engine's synthesis API with adjustable voice, speed, and tone. Byted-Text-to-Speech transforms written content into audio using Volcano Engine's speech synthesis service. Configure speaker voice, speech rate, pitch, volume, and output format—with optional Markdown filtering and LaTeX formula support for technical content.

manual: git clone https://github.com/bytedance/agentkit-samples → cp -r agentkit-samples/skills/byted-text-to-speech ~/.claude/skills/byted-text-to-speech
skills/byted-text-to-speech/SKILL.md · version 42e0dadf

Use it when

  • Yes.
  • Byted-text-to-speech converts written text into spoken audio by processing your input through Volcano Engine's synthesis engine.

Verify before relying

Read SKILL.md below before installing (7 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

bytedance/agentkit-samples/byted-text-to-speech · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What does byted-text-to-speech do?

Byted-text-to-speech transforms written content into audio using Volcano Engine's speech synthesis service. You can configure speaker voice, speech rate, pitch, volume, and output format—with optional Markdown filtering and LaTeX formula support for technical content.

Can I convert text to speech with different voices?

Yes. Byted-text-to-speech supports multiple voice options and lets you customize speech synthesis with adjustable speed, pitch, volume, and different speaker selections to suit your narration or voiceover needs.

How do I read text aloud using this tool?

Byted-text-to-speech converts written text into spoken audio by processing your input through Volcano Engine's synthesis engine. Simply provide your text, select your preferred voice and pacing settings, and the skill generates audio output in your chosen format.

Does byted-text-to-speech support markdown and LaTeX?

Yes. Byted-text-to-speech includes optional Markdown filtering and LaTeX formula support, making it suitable for converting technical documents, academic content, and formatted text into natural-sounding audio.

Can I generate podcast audio or voiceovers with this skill?

Yes. Byted-text-to-speech is designed to generate voiceovers, narration, and audiobook segments from text content. You can create podcast audio by configuring voice, speed, and other synthesis parameters to match your production needs.

What languages does byted-text-to-speech support?

Byted-text-to-speech creates multilingual audio output from text, allowing you to synthesize speech in multiple languages. Combined with its Markdown and LaTeX support, it handles diverse content types across language boundaries.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

Byted-Text-to-Speech Skill

基于火山引擎豆包语音合成(HTTP Chunked/SSE 单向流式-V3)将文本转为语音并保存为音频文件。

何时使用

当用户有以下需求时,优先使用本 skill:

  • 需要把一段文字转成语音、朗读音频
  • 需要生成配音、旁白、播报、有声读物片段
  • 需要将代码注释、文档、文章等内容转为音频便于收听
  • 需要生成多语言语音(中文、英文等)
  • 用户提到「文字转语音」「TTS」「语音合成」「朗读」「配音」「念出来」「读给我听」
  • 用户没有明确说"语音合成",但任务本质上需要将文本内容转为可播放的音频时

使用前检查

优先检查是否已配置以下凭证:

  • MODEL_SPEECH_API_KEY

如果缺少凭证,打开 references/setup-guide.md 查看开通、申请和配置方式,并给予用户开通建议

脚本参数

参数 简写 必填 说明
--text -t 要合成的文本内容
--output -o 输出音频文件路径(默认自动生成)
--speaker -s 发音人,默认 zh_female_vv_uranus_bigtts音色列表
--format 音频格式:mp3(默认)、pcmogg_opus
--sample-rate 采样率,如 16000、24000(默认 24000)
--speech-rate 语速 [-50, 100],100 代表

(truncated - see the full file via the links below)

File tree — 7 files
skills/byted-text-to-speech/LICENSE
skills/byted-text-to-speech/SKILL.md
skills/byted-text-to-speech/references/docs-index.md
skills/byted-text-to-speech/references/setup-guide.md
skills/byted-text-to-speech/requirements.txt
skills/byted-text-to-speech/scripts/api_key.py
skills/byted-text-to-speech/scripts/text_to_speech.py

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert written text into spoken audio with customizable voice and pacing”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

byted-podcast-gen
by bytedance · bytedance/agentkit-samples

byted-podcast-gen synthesizes podcast audio from a topic phrase, web URL, or long-form text using Volcano Engine's PodcastTTS protocol. The skill handles three input modes—topic-based, URL-based, and file-based—and outputs audio files in MP3, WAV, or OGG format along with segmented speaker text.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
byted-voice-to-text
by bytedance · bytedance/agentkit-samples

Byted Voice to Text transcribes audio using Volcano Engine's BigModel ASR, offering fast synchronous processing for files under 2 hours and 100MB, or asynchronous recognition for longer content up to 5 hours. It handles Feishu voice messages, local audio files, and direct URLs with automatic format detection.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
podcast-generator
by staruhub · staruhub/ClaudeSkills

Transform Chinese articles and reports into conversational podcast audio featuring two speakers. The skill handles format selection (MP3, OGG Opus, PCM, AAC), speech rate adjustment, voice customization, and resume-on-failure for interrupted generations. Requires Volcano Engine credentials and works best with structured text between 500–3000 characters.

MITdocs in Chineseupdated Jul 2026
★ 631repo stars
listenhub-tts
by smallnest · smallnest/goal-workflow

ListenHub TTS transforms written content into spoken audio through three synthesis modes: rapid processing for short text, multi-speaker dialogue for scripts, and streaming synthesis for lengthy documents. Select from available voices or use the default voice, with options to adjust playback speed and output format.

MITdocs in Chineseupdated Jul 2026
★ 181repo stars
byted-web-search
by bytedance · bytedance/agentkit-samples

Byted Web Search connects your agent to live internet data via Volcengine's official API, delivering web and image results for fact-checking, current events, and time-sensitive queries. Designed for scenarios where knowledge cutoff limits accuracy, it triggers automatically on temporal keywords and verification requests, with 500 free monthly searches included.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
byted-las-audio-convert
by bytedance · bytedance/agentkit-samples

This skill transcodes audio files across multiple formats and adjusts encoding parameters like sample rate, bitrate, and channels through Volcengine's LAS service. It accepts files from TOS cloud storage or local uploads, processes them via ffmpeg-backed parameters, and returns converted audio ready for downstream use.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars

More skills Funasr Transcribe (unlicensed) · byted-vms-voice-notify (Apache-2.0)

Tags
voice-synthesisaudio-generationnarration-enginespeech-customizationmultilingual-supportstreaming-synthesisvoice-cloningaccessibility-toolcontent-to-audio