$npx skillfedfor your agent

mimo-v2-5-tts

MiMo V2.5 TTS transforms text into natural speech across three modes: preset voices for quick synthesis, voice design for custom tones via text description, and voice cloning from audio samples. Control emotion, dialect, and style through natural language, audio tags, or director mode for cinematic-quality output.

MiMo V2.5 TTS converts text to speech with preset voices, custom voice design, or audio cloning.

AI-generated summary based on this skill's SKILL.md

86 11 MITupdated by XiaomiMiMo

Decision gist · record as of 2026-04-24

MiMo V2.5 TTS converts text to speech with preset voices, custom voice design, or audio cloning. MiMo V2.5 TTS transforms text into natural speech across three modes: preset voices for quick synthesis, voice design for custom tones via text description, and voice cloning from audio samples. Control emotion, dialect, and style through natural language, audio tags, or director mode for cinematic-quality output.

manual: git clone https://github.com/XiaomiMiMo/MiMo-Skills → cp -r MiMo-Skills/skills/mimo-v2-5-tts ~/.claude/skills/mimo-v2-5-tts
skills/mimo-v2-5-tts/SKILL.md · version f72ed60f

Use it when

  • MiMo V2.5 TTS enables voice cloning by analyzing audio samples you provide.
  • Yes, MiMo V2.5 TTS provides fine-grained emotional and stylistic control over generated speech.

Verify before relying

Read SKILL.md below before installing (5 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

XiaomiMiMo/MiMo-Skills/mimo-v2-5-tts · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What can MiMo V2.5 TTS do with text to speech Chinese?

MiMo V2.5 TTS converts Chinese text into natural speech with multiple customization options. You can choose from preset voices for immediate synthesis, design custom voices by describing the tone you want, or clone voices from audio samples. The skill supports emotional expression, dialect variations, and stylistic control through natural language commands, audio tags, or director mode for professional-quality output.

How does voice cloning work in MiMo V2.5 TTS?

MiMo V2.5 TTS enables voice cloning by analyzing audio samples you provide, then generating speech that matches that voice's characteristics. Beyond simple cloning, you can also design entirely custom voices by providing text descriptions of the tone, accent, or personality you want. This flexibility lets you create unique voice personas tailored to your specific needs.

Can MiMo V2.5 TTS generate speech with specific emotions?

Yes, MiMo V2.5 TTS provides fine-grained emotional and stylistic control over generated speech. You can direct the emotion, tone, and delivery style through natural language instructions, audio tags, or director mode—a feature designed for cinematic-quality voice acting. This allows expressive synthesis that goes beyond neutral reading.

Does MiMo V2.5 TTS support singing or dialect-specific speech?

MiMo V2.5 TTS supports both singing voice synthesis and dialect-specific speech generation. Whether you need content in different regional accents or want to synthesize singing, the skill handles these specialized audio tasks. Combined with its emotional control features, you can create diverse vocal content across multiple languages and styles.

What are the three main modes for generating speech in MiMo V2.5 TTS?

MiMo V2.5 TTS operates in three core modes: preset voices for quick, ready-to-use synthesis; voice design for creating custom tones via text description; and voice cloning from audio samples. Each mode supports emotional and stylistic customization, letting you choose between convenience and personalization depending on your project needs.

Can MiMo V2.5 TTS send voice messages to Feishu chat?

MiMo V2.5 TTS can generate voice messages and send them to Feishu chat or private conversations. This integration enables you to deliver synthesized speech content directly within your messaging workflow, combining the skill's text-to-speech capabilities with seamless team communication.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

MiMo V2.5 TTS

使用小米 MiMo V2.5 TTS 系列模型生成语音。支持中英文、预置音色、音色设计、音色克隆、情绪风格、方言、唱歌。

脚本目录:$SKILLS_PATH/mimo-v2-5-tts/scripts/

> $SKILLS_PATH 说明: skills 目录路径,因部署环境而异。

模型选择

V2.5 系列提供三种模型,根据使用场景选择:

模型 ID 用途 音色来源 特殊能力
mimo-v2.5-tts 预置音色语音合成 内置精品音色 支持唱歌
mimo-v2.5-tts-voicedesign 文本描述定制音色 文本描述生成
mimo-v2.5-tts-voiceclone 音频样本复刻音色 音频样本

选择建议:

  • 需要快速生成语音、需要唱歌功能 → mimo-v2.5-tts(预置音色)
  • 需要独特音色 → mimo-v2.5-tts-voicedesign(文本描述生成)
  • 需要模仿特定声音 → mimo-v2.5-tts-voiceclone(音频样本复刻)

> 注意: TTS 有随机性,同样输入的效果可能不同,用户有需要时可以多生成几次以供挑选。

环境依赖

环境变量 说明 必需
MIMO_API_KEY MiMo API 密钥(MiMo 开放平台获取)

| 依赖 | 说明 | 必需

(truncated - see the full file via the links below)

File tree — 5 files
skills/mimo-v2-5-tts/SKILL.md
skills/mimo-v2-5-tts/scripts/feishu_send_audio.sh
skills/mimo-v2-5-tts/scripts/mimo_tts.py
skills/mimo-v2-5-tts/scripts/mimo_tts_voiceclone.py
skills/mimo-v2-5-tts/scripts/mimo_tts_voicedesign.py

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert text to natural speech with customizable voice and emotion”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

feishu-send-file
by Rabbitmeaw · Rabbitmeaw/feishu-send-file

This skill delivers files, images, and voice messages to Feishu users through the platform's required two-step workflow. Upload your content first to obtain a file or image key, then send the message using that key. Supports documents, pictures, and opus audio formats.

MITdocs in Chineseupdated Mar 2026
★ 16repo stars
byted-vms-voice-notify
by bytedance · bytedance/agentkit-samples

Byted-vms-voice-notify wraps Volcano Engine's voice notification API to deliver single or batch voice calls to phone numbers. It handles TTS template lifecycle (create, update, delete, query), recording file uploads and management, and resource operations—serving as the unified gateway for all voice notification and template operations.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
Feishu Voice Tts
by Tangc · Tangc/tangzhan-skills

Feishu Voice TTS transforms text into speech using edge-tts and delivers it as Feishu audio messages, bypassing the platform's text fallback for direct file sends. The skill handles transcoding to Opus format, file upload, and message delivery through Feishu's open API.

no license declared → metadata onlyupdated Jun 2026
★ 5repo stars
listenhub-tts
by smallnest · smallnest/goal-workflow

ListenHub TTS transforms written content into spoken audio through three synthesis modes: rapid processing for short text, multi-speaker dialogue for scripts, and streaming synthesis for lengthy documents. Select from available voices or use the default voice, with options to adjust playback speed and output format.

MITdocs in Chineseupdated Jul 2026
★ 181repo stars
byted-voice-to-text
by bytedance · bytedance/agentkit-samples

Byted Voice to Text transcribes audio using Volcano Engine's BigModel ASR, offering fast synchronous processing for files under 2 hours and 100MB, or asynchronous recognition for longer content up to 5 hours. It handles Feishu voice messages, local audio files, and direct URLs with automatic format detection.

Apache-2.0docs in Chineseupdated Jul 2026
★ 378repo stars
Ttscn
by Agents365-ai · Agents365-ai/365-skills

Ttscn converts Chinese and multilingual text to natural speech across 11 cloud backends, from free Edge TTS to premium providers like ElevenLabs and OpenAI. Choose by use case—short video, audiobook, enterprise SSML, or lowest cost—with support for word-level timestamps, pause markers, and pronunciation overrides.

no license declared → metadata onlyupdated Jul 2026
★ 24repo stars

More skills Funasr Transcribe (unlicensed)

Tags
voice-cloningemotional-synthesismultilingual-ttsvoice-designcharacter-voicesdialect-supportsinging-synthesisdirector-modeaudio-taggingpreset-voices