qwencloud-audio-tts
Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.
qwencloud-audio-tts converts text to natural speech audio using Qwen TTS models via HTTP or WebSocket APIs.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-07-22
qwencloud-audio-tts converts text to natural speech audio using Qwen TTS models via HTTP or WebSocket APIs. Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.
Use it when
- qwencloud-audio-tts provides multiple TTS models for text-to-audio conversion.
- Yes, qwencloud-audio-tts is designed for generating voiceovers and audio narration for content.
Verify before relying
Read SKILL.md below before installing (11 files). Open directory: indexed for reading, not audited.
Install
QwenCloud/qwencloud-ai/qwencloud-audio-tts · repository language: Python
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What does qwencloud-audio-tts do?
qwencloud-audio-tts converts written text into high-quality spoken audio using QwenCloud's TTS models. You can choose between fast standard synthesis with qwen3-tts-flash, instruction-guided style control with qwen3-tts-instruct-flash, or premium quality via CosyVoice. The skill supports multiple voices and languages to match your content needs.
How can I convert text to audio with qwencloud-audio-tts?
qwencloud-audio-tts provides multiple TTS models for text-to-audio conversion. Use qwen3-tts-flash for fast, standard synthesis, or qwen3-tts-instruct-flash when you need instruction-guided style control over the output. For premium quality audio, CosyVoice is available. Select your preferred voice and language, then submit your text to generate natural speech audio.
Can I generate voiceovers and narration using qwencloud-audio-tts?
Yes, qwencloud-audio-tts is designed for generating voiceovers and audio narration for content. The instruction-guided qwen3-tts-instruct-flash model lets you control speech style, while the standard qwen3-tts-flash model offers fast synthesis. You can select from multiple voices and languages to create professional narration that matches your content.
What TTS voice generation options does qwencloud-audio-tts offer?
qwencloud-audio-tts provides multiple voice and language options across three model tiers. The qwen3-tts-flash model delivers fast, standard voice synthesis. The qwen3-tts-instruct-flash model adds instruction-guided style control for customized speech output. CosyVoice offers premium quality audio. Choose the model and voice that best fits your application's performance and quality requirements.
How do I build or integrate text-to-speech into my application?
qwencloud-audio-tts provides a text-to-speech API for application integration. Select your preferred TTS model—qwen3-tts-flash for speed, qwen3-tts-instruct-flash for style control, or CosyVoice for premium quality—then integrate the API into your application workflow. The skill supports multiple voices and languages, enabling flexible voice narration generation across diverse use cases.
What speech synthesis models are available in qwencloud-audio-tts?
qwencloud-audio-tts offers three speech synthesis options: qwen3-tts-flash for fast, standard audio output; qwen3-tts-instruct-flash for instruction-guided style control over the generated speech; and CosyVoice for premium-quality audio. All models support multiple voices and languages, letting you synthesize natural speech audio tailored to your content and performance needs.
SKILL.md
Rendered from the published skill. Quoted content, verbatim.
> Agent setup: If your agent doesn't auto-load skills (e.g. Claude Code), > see agent-compatibility.md once per session.
Qwen Audio TTS (Text-to-Speech)
Synthesize natural speech from text using Qwen TTS models. This skill is part of qwencloud/qwencloud-ai.
Skill directory
Use this skill's internal files to execute and learn. Load reference files on demand when the default path fails or you need details.
| Location | Purpose |
|---|---|
scripts/tts.py |
Qwen TTS (HTTP API) — qwen3-tts-flash, |
(truncated - see the full file via the links below)
File tree — 11 files
skills/audio/qwencloud-audio-tts/SKILL.md
skills/audio/qwencloud-audio-tts/references/agent-compatibility.md
skills/audio/qwencloud-audio-tts/references/api-guide.md
skills/audio/qwencloud-audio-tts/references/cosyvoice-guide.md
skills/audio/qwencloud-audio-tts/references/execution-guide.md
skills/audio/qwencloud-audio-tts/references/prompt-guide.md
skills/audio/qwencloud-audio-tts/references/sources.md
skills/audio/qwencloud-audio-tts/scripts/gossamer.py
skills/audio/qwencloud-audio-tts/scripts/qwencloud_lib.py
skills/audio/qwencloud-audio-tts/scripts/tts.py
skills/audio/qwencloud-audio-tts/scripts/tts_cosyvoice.py
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Convert text to natural speech audio using Qwen TTS models”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
Qwen Audio TTS turns written text into natural-sounding speech using Qwen's TTS engine. Choose from multiple voices and models—including the fast qwen3-tts-flash for standard tasks or instruction-guided variants for tone control—then output audio directly to file.
Access Qwen's language models for text generation, multi-turn conversations, code writing, and function calling through an OpenAI-compatible interface. The skill supports multiple Qwen variants optimized for different tasks—from general-purpose models to specialized code and reasoning versions—with flexible model selection and streaming output.
Ttscn converts Chinese and multilingual text to natural speech across 11 cloud backends, from free Edge TTS to premium providers like ElevenLabs and OpenAI. Choose by use case—short video, audiobook, enterprise SSML, or lowest cost—with support for word-level timestamps, pause markers, and pronunciation overrides.
Qwen Vision lets you understand images and videos through Qwen's specialized VL and QVQ models. Extract text via OCR, analyze charts and tables, perform visual reasoning, and compare multiple images—all with built-in support for thinking mode and high-resolution processing.
Create images from text descriptions or edit existing ones with Wan and Qwen Image models. This skill handles text-to-image generation, style transfer, subject consistency across reference images, and interleaved text-image output for tutorials and guides.
Tts transforms written content into spoken audio through two synthesis backends—Kokoro for local processing and Noiz for advanced features like voice cloning and emotion mapping. Use simple mode for quick narration or timeline mode to align speech precisely to subtitle segments for video dubbing and audiobook production.
More skills qwencloud-video-generation (Apache-2.0) · Dashscope (AGPL-3.0) · qianwen-text (Apache-2.0) · Text To Speech (Unlicense) · qianwen-image-generation (Apache-2.0) · aliyun-modelstudio-entry-test (MIT)