$npx skillfedfor your agent

Doubao Tts

Doubao TTS produces natural Mandarin and multilingual audio narration via Volcengine's Speech 2.0 API. It returns word-level timing metadata for precise subtitle synchronization, making it ideal for video projects requiring accurate caption alignment. Configure your voice preference and speech rate, then generate samples before committing to full narrations.

Doubao TTS converts text to Mandarin and multilingual speech with character-level timing for subtitles.

AI-generated summary based on this skill's SKILL.md

42,764 5,160 AGPL-3.0updated by calesthio

Decision gist · record as of 2026-07-24

Doubao TTS converts text to Mandarin and multilingual speech with character-level timing for subtitles. Doubao TTS produces natural Mandarin and multilingual audio narration via Volcengine's Speech 2.0 API. It returns word-level timing metadata for precise subtitle synchronization, making it ideal for video projects requiring accurate caption alignment. Configure your voice preference and speech rate, then generate samples before committing to full narrations.

manual: git clone https://github.com/calesthio/OpenMontage → cp -r OpenMontage ~/.claude/skills/doubao-tts

Use it when

  • Yes.
  • Doubao TTS returns word-level timing metadata alongside the generated audio.
Same gist for agents: .md · .json

Install

calesthio/OpenMontage/doubao-tts · repository language: Python

generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What does Doubao TTS do?

Doubao TTS produces natural Mandarin and multilingual audio narration via Volcengine's Speech 2.0 API. It generates spoken audio from written text and returns word-level timing metadata for precise subtitle synchronization, making it ideal for video projects requiring accurate caption alignment.

Can Doubao TTS convert text to speech?

Yes. Doubao TTS converts text to speech using Doubao's voice synthesis capability. You can configure your voice preference and speech rate, then generate samples before committing to full narrations. The skill produces natural audio output across Mandarin and multiple languages.

How does Doubao voice synthesis work for subtitles?

Doubao TTS returns word-level timing metadata alongside the generated audio. This precise timing information enables accurate subtitle synchronization, allowing you to align captions exactly with the spoken content for video projects and multimedia applications.

What customization options does Doubao TTS offer?

Doubao TTS lets you configure voice preference and speech rate to match your project needs. You can generate audio samples with your chosen settings before processing full narrations, ensuring the output meets your quality and style requirements.

Which languages does Doubao TTS support?

Doubao TTS produces natural audio narration in Mandarin and multiple other languages via Volcengine's Speech 2.0 API, making it suitable for diverse multilingual content creation and global video projects.

What is the license for Doubao TTS?

Doubao TTS is licensed under AGPL-3.0, which requires that any modifications or derivative works remain open source and available under the same license terms.

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert text to speech using Doubao TTS”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

Dashscope
by calesthio · calesthio/OpenMontage

Connect Dashscope to harness Alibaba Cloud's Qwen models for multimodal content creation. Generate images via qwen-image-2.0-pro, synthesize speech with qwen3-tts-flash, or transcribe audio with word-level timing via qwen3-asr-flash-filetrans—all through native DashScope endpoints.

AGPL-3.0updated Jul 2026
★ 42,764repo stars
Ttscn
by Agents365-ai · Agents365-ai/365-skills

Ttscn converts Chinese and multilingual text to natural speech across 11 cloud backends, from free Edge TTS to premium providers like ElevenLabs and OpenAI. Choose by use case—short video, audiobook, enterprise SSML, or lowest cost—with support for word-level timestamps, pause markers, and pronunciation overrides.

no license declared → metadata onlyupdated Jul 2026
★ 24repo stars
Article Explainer Video
by wwwzhouhui · wwwzhouhui/skills_collection

Article Explainer Video transforms long-form technical content into 5–8 minute narrated videos with multiple layout styles, AI-generated illustrations, and synchronized subtitles. Choose between warm (cozy handdrawn) or midnight (tech-forward) visual themes, then let the pipeline handle storyboarding, illustration generation, text-to-speech, and MP4 rendering.

no license declared → metadata onlyupdated Jul 2026
★ 256repo stars
Kling Official
by calesthio · calesthio/OpenMontage

Kling Official provides direct API guidance for video, image, TTS, avatar, and lip-sync generation through Kling's native endpoints. It handles authentication via KLING_API_KEY, manages both classic and turbo task protocols, and keeps operations separate from fal.ai routing.

AGPL-3.0updated Jul 2026
★ 42,764repo stars
Tts
by NoizAI · NoizAI/skills

Tts transforms written content into spoken audio through two synthesis backends—Kokoro for local processing and Noiz for advanced features like voice cloning and emotion mapping. Use simple mode for quick narration or timeline mode to align speech precisely to subtitle segments for video dubbing and audiobook production.

no license declared → metadata onlyupdated May 2026
★ 524repo stars
voice-localization
by guia-matthieu · guia-matthieu/clawfu-skills

This skill guides you through scaling video and audio content to global audiences using AI voice synthesis that preserves your brand character across languages. It provides decision frameworks for choosing between AI localization, traditional dubbing, and subtitles based on your content type and budget, plus production workflows that handle translation, voice generation, and quality assurance per market. Use it to expand into new language markets efficiently while keeping the same perceived voice speaking natively in each language.

MITupdated Apr 2026
★ 140repo stars

More skills doubao-tts (MIT)

Tags
voice-synthesisaudio-generationspeech-outputtext-conversionnarration-enginevoice-apiaudio-servicespeech-processing