$npx skillfedfor your agent

listenhub-tts

ListenHub TTS transforms written content into spoken audio through three synthesis modes: rapid processing for short text, multi-speaker dialogue for scripts, and streaming synthesis for lengthy documents. Select from available voices or use the default voice, with options to adjust playback speed and output format.

ListenHub TTS converts text to speech audio via API with support for quick synthesis, multi-character scripts, and long-form streaming.

AI-generated summary based on this skill's SKILL.md

181 28 MITupdated by smallnest

Decision gist · record as of 2026-07-21

ListenHub TTS converts text to speech audio via API with support for quick synthesis, multi-character scripts, and long-form streaming. ListenHub TTS transforms written content into spoken audio through three synthesis modes: rapid processing for short text, multi-speaker dialogue for scripts, and streaming synthesis for lengthy documents. Select from available voices or use the default voice, with options to adjust playback speed and output format.

manual: git clone https://github.com/smallnest/goal-workflow → cp -r goal-workflow/skills/listenhub-tts ~/.claude/skills/listenhub-tts
skills/listenhub-tts/SKILL.md · version 19552be3

Use it when

  • Yes, ListenHub TTS supports multi-speaker dialogue synthesis, making it suitable for podcast scripts and character-driven content.
  • ListenHub TTS includes support for Chinese text-to-speech synthesis.

Verify before relying

Read SKILL.md below before installing (1 file). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

smallnest/goal-workflow/listenhub-tts · repository language: HTML

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

How does ListenHub TTS convert text to speech audio?

ListenHub TTS transforms written content into spoken audio through three synthesis modes: rapid processing for short text, multi-speaker dialogue for scripts, and streaming synthesis for lengthy documents. You can select from available voices or use the default voice, with options to adjust playback speed and output format.

Can ListenHub TTS generate multi-character dialogue or podcast narration?

Yes, ListenHub TTS supports multi-speaker dialogue synthesis, making it suitable for podcast scripts and character-driven content. The skill allows you to assign different voices to different speakers, enabling natural-sounding conversations and multi-character narration for your audio projects.

Does ListenHub TTS support text to audio synthesis in Chinese?

ListenHub TTS includes support for Chinese text-to-speech synthesis. The skill can process Chinese text and generate audio output, making it suitable for Chinese language content, articles, and documents that need to be converted to spoken audio.

How does ListenHub TTS handle long-form content?

ListenHub TTS provides streaming synthesis support for lengthy documents and long-form content. This enables efficient processing of extended texts such as articles, books, or comprehensive documents without requiring the entire content to be processed at once.

What voice customization options does ListenHub TTS offer?

ListenHub TTS allows you to select and apply custom voice speakers for your text-to-speech synthesis. Beyond voice selection, you can adjust playback speed and choose your preferred output format, giving you control over how your audio content sounds and is delivered.

Can ListenHub TTS automate article or document narration?

Yes, ListenHub TTS can automate the narration of articles and documents by converting written text directly into audio. This is useful for creating audiobook versions, generating voice-over content, or making written material more accessible through automated speech synthesis.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

ListenHub TTS: 文本转语音

使用 ListenHub OpenAPI 将文本转换为语音。支持三种合成模式,覆盖从短文本到长文本的全场景。

API 信息

  • Base URL: https://api.marswave.ai/openapi
  • 认证: Authorization: Bearer $LISTENHUB_API_KEY(从环境变量读取)
  • 前置检查: 调用任何 API 前先确认 LISTENHUB_API_KEY 环境变量已设置,未设置则提示用户配置

音色选择流程

用户已明确指定音色

直接使用用户指定的 speakerId,跳过选择流程。

用户未指定音色
  1. 调用 GET /v1/speakers/list?language=zh 获取可用音色列表
  2. 按 AskUserQuestion 展示音色列表供用户选择,格式如下:
  3. 默认选中 chat-girl-105-cn(晓曼 dxqqq)
  4. 列表展示:{name}({gender},{speakerId})
  5. 附带每个音色的 demoAudioUrl 供参考
  6. 用户确认后使用选定的 speakerId
默认音色
字段
speakerId chat-girl-105-cn
名称 晓曼 dxqqq

三种合成模式

模式一:快速合成(短文本,单音色)

适用场景: 短文本(< 1000 字),单音色,需要低延迟

接口:

(truncated - see the full file via the links below)

File tree — 1 file
skills/listenhub-tts/SKILL.md

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert text to speech audio using ListenHub API”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

podcast-generator
by staruhub · staruhub/ClaudeSkills

Transform Chinese articles and reports into conversational podcast audio featuring two speakers. The skill handles format selection (MP3, OGG Opus, PCM, AAC), speech rate adjustment, voice customization, and resume-on-failure for interrupted generations. Requires Volcano Engine credentials and works best with structured text between 500–3000 characters.

MITdocs in Chineseupdated Jul 2026
★ 631repo stars
byted-text-to-speech
by bytedance · bytedance/agentkit-samples

Byted-Text-to-Speech transforms written content into audio using Volcano Engine's speech synthesis service. Configure speaker voice, speech rate, pitch, volume, and output format—with optional Markdown filtering and LaTeX formula support for technical content.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
doubao-tts
by xvirobotics · xvirobotics/metabot

Doubao TTS converts text to natural-sounding speech via Volcengine's API, handling both quick responses under 300 characters and longer content up to 100K characters asynchronously. Choose from multiple Chinese male and female voices, adjust speed and pitch, and build podcasts with multi-voice support.

MITfor claude-codeupdated Jul 2026
★ 934repo stars
humanize-it
by smallnest · smallnest/goal-workflow

Humanize-It automatically detects your document type and applies the most effective de-AI rewriting strategy, cycling through specialized skills until the text reads naturally human-written. It handles general articles, technical documentation, and academic content with up to 42 iterative passes.

MITdocs in Chineseupdated Jul 2026
★ 181repo stars
Text To Speech
by martinholovsky · martinholovsky/claude-skills-generator

This skill provides expert-level text-to-speech implementation using Kokoro TTS, enabling real-time voice synthesis with customizable voices and prosody control. It emphasizes secure content handling, performance optimization through streaming and caching, and resource-efficient audio generation suitable for voice assistant applications.

Unlicenseupdated Dec 2025
★ 45repo stars
Ttscn
by Agents365-ai · Agents365-ai/365-skills

Ttscn converts Chinese and multilingual text to natural speech across 11 cloud backends, from free Edge TTS to premium providers like ElevenLabs and OpenAI. Choose by use case—short video, audiobook, enterprise SSML, or lowest cost—with support for word-level timestamps, pause markers, and pronunciation overrides.

no license declared → metadata onlyupdated Jul 2026
★ 24repo stars

More skills mimo-v2-5-tts (MIT) · dialogue-manager (MIT) · Feishu Voice Tts (unlicensed) · Tts (unlicensed)

Tags
audio-synthesisvoice-generationmulti-speakerstreaming-synthesischinese-languagenarration-automationpodcast-productionapi-integrationlong-form-contentspeaker-selection