Tts
Tts transforms written content into spoken audio through two synthesis backends—Kokoro for local processing and Noiz for advanced features like voice cloning and emotion mapping. Use simple mode for quick narration or timeline mode to align speech precisely to subtitle segments for video dubbing and audiobook production.
Tts converts written text into spoken audio with support for voice cloning and precise timeline alignment.
AI-generated summary based on this skill's SKILL.md
Decision gist · record as of 2026-05-07
Tts converts written text into spoken audio with support for voice cloning and precise timeline alignment. Tts transforms written content into spoken audio through two synthesis backends—Kokoro for local processing and Noiz for advanced features like voice cloning and emotion mapping. Use simple mode for quick narration or timeline mode to align speech precisely to subtitle segments for video dubbing and audiobook production.
Use it when
- Yes, Tts converts text to audio using two backends.
- Tts offers simple mode for quick narration tasks and timeline mode for precise alignment of speech to subtitle segments.
Similar skills
Install
NoizAI/skills/tts · repository language: Python
generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub
Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.
Frequently asked questions
AI-generated answers based on this skill's SKILL.md and metadata
What does Tts do?
Tts transforms written content into spoken audio through two synthesis backends—Kokoro for local processing and Noiz for advanced features like voice cloning and emotion mapping. Use simple mode for quick narration or timeline mode to align speech precisely to subtitle segments for video dubbing and audiobook production.
Can Tts convert text to audio?
Yes, Tts converts text to audio using two backends. Kokoro handles local processing for straightforward narration, while Noiz provides advanced capabilities including voice cloning and emotion mapping to create more expressive and personalized audio output from your written text.
What are Tts's two main modes?
Tts offers simple mode for quick narration tasks and timeline mode for precise alignment of speech to subtitle segments. Timeline mode is particularly useful for video dubbing and audiobook production where you need synchronized audio matching specific text timing.
Does Tts support voice cloning and emotion mapping?
Yes, Tts's Noiz backend supports voice cloning and emotion mapping features. These advanced capabilities allow you to generate natural-sounding voice narration with personalized vocal characteristics and emotional expression for more engaging audio content.
How can Tts help with accessibility?
Tts creates accessibility-friendly audio content by converting written text into spoken audio, making content available to users who prefer or require audio formats. This automation also enables efficient voice-over production for videos, audiobooks, and other media without manual recording.
Let your AI agent find skills like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.
wish › “Convert written text into spoken audio”
Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →
Related skills
This skill provides expert-level text-to-speech implementation using Kokoro TTS, enabling real-time voice synthesis with customizable voices and prosody control. It emphasizes secure content handling, performance optimization through streaming and caching, and resource-efficient audio generation suitable for voice assistant applications.
Ttscn converts Chinese and multilingual text to natural speech across 11 cloud backends, from free Edge TTS to premium providers like ElevenLabs and OpenAI. Choose by use case—short video, audiobook, enterprise SSML, or lowest cost—with support for word-level timestamps, pause markers, and pronunciation overrides.
9router-tts routes text-to-speech requests to your choice of seven major providers—OpenAI, ElevenLabs, Deepgram, Edge TTS, Google TTS, Hyperbolic, and Inworld—through a unified API endpoint. Query available models and voices per provider, then POST your text with a voice ID to receive MP3 audio or base64-encoded JSON. Each provider has its own authentication and voice naming scheme, all abstracted behind a single interface.
Power Iterate is a fully autonomous skill that handles repetitive development work without interruption, managing time and token budgets while executing iteratively toward completion. It understands requirements automatically, designs evaluation criteria, and executes tasks in cycles—stopping only when budgets run out or quality gates are met. Ideal for sustained coding sessions where you set parameters and let the skill drive progress.
Doubao TTS produces natural Mandarin and multilingual audio narration via Volcengine's Speech 2.0 API. It returns word-level timing metadata for precise subtitle synchronization, making it ideal for video projects requiring accurate caption alignment. Configure your voice preference and speech rate, then generate samples before committing to full narrations.
Turn written text into high-quality spoken audio using QwenCloud's TTS models. Choose between fast standard synthesis (qwen3-tts-flash) or instruction-guided style control (qwen3-tts-instruct-flash), or opt for premium quality via CosyVoice. Select from multiple voices and languages to match your content needs.
More skills qianwen-audio-tts (Apache-2.0) · happy-audio-gen (MIT) · listenhub-tts (MIT)