skillfed

mimo-v2-5-tts

MiMo V2.5 TTS transforms text into natural speech across three modes: preset voices for quick synthesis, voice design for custom tones via text description, and voice cloning from audio samples. Control emotion, dialect, and style through natural language, audio tags, or director mode for cinematic-quality output.

MiMo V2.5 TTS converts text to speech with preset voices, custom voice design, or audio cloning.

AI-generated summary based on this skill's SKILL.md

86 11 MIT updated by XiaomiMiMo

Install

XiaomiMiMo/MiMo-Skills/mimo-v2-5-tts · repository language: Python

CLI (skillfed)coming soon
git clone https://github.com/XiaomiMiMo/MiMo-Skills
cp -r MiMo-Skills/skills/mimo-v2-5-tts ~/.claude/skills/mimo-v2-5-tts

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What can MiMo V2.5 TTS do with text to speech Chinese?

MiMo V2.5 TTS converts Chinese text into natural speech with multiple customization options. You can choose from preset voices for immediate synthesis, design custom voices by describing the tone you want, or clone voices from audio samples. The skill supports emotional expression, dialect variations, and stylistic control through natural language commands, audio tags, or director mode for professional-quality output.

How does voice cloning work in MiMo V2.5 TTS?

MiMo V2.5 TTS enables voice cloning by analyzing audio samples you provide, then generating speech that matches that voice's characteristics. Beyond simple cloning, you can also design entirely custom voices by providing text descriptions of the tone, accent, or personality you want. This flexibility lets you create unique voice personas tailored to your specific needs.

Can MiMo V2.5 TTS generate speech with specific emotions?

Yes, MiMo V2.5 TTS provides fine-grained emotional and stylistic control over generated speech. You can direct the emotion, tone, and delivery style through natural language instructions, audio tags, or director mode—a feature designed for cinematic-quality voice acting. This allows expressive synthesis that goes beyond neutral reading.

Does MiMo V2.5 TTS support singing or dialect-specific speech?

MiMo V2.5 TTS supports both singing voice synthesis and dialect-specific speech generation. Whether you need content in different regional accents or want to synthesize singing, the skill handles these specialized audio tasks. Combined with its emotional control features, you can create diverse vocal content across multiple languages and styles.

What are the three main modes for generating speech in MiMo V2.5 TTS?

MiMo V2.5 TTS operates in three core modes: preset voices for quick, ready-to-use synthesis; voice design for creating custom tones via text description; and voice cloning from audio samples. Each mode supports emotional and stylistic customization, letting you choose between convenience and personalization depending on your project needs.

Can MiMo V2.5 TTS send voice messages to Feishu chat?

MiMo V2.5 TTS can generate voice messages and send them to Feishu chat or private conversations. This integration enables you to deliver synthesized speech content directly within your messaging workflow, combining the skill's text-to-speech capabilities with seamless team communication.

SKILL.md

rendered from the published skill — quoted content, verbatim

MiMo V2.5 TTS

使用小米 MiMo V2.5 TTS 系列模型生成语音。支持中英文、预置音色、音色设计、音色克隆、情绪风格、方言、唱歌。

脚本目录:$SKILLS_PATH/mimo-v2-5-tts/scripts/

> $SKILLS_PATH 说明: skills 目录路径,因部署环境而异。

模型选择

V2.5 系列提供三种模型,根据使用场景选择:

模型 ID 用途 音色来源 特殊能力
mimo-v2.5-tts 预置音色语音合成 内置精品音色 支持唱歌
mimo-v2.5-tts-voicedesign 文本描述定制音色 文本描述生成
mimo-v2.5-tts-voiceclone 音频样本复刻音色 音频样本

选择建议:

  • 需要快速生成语音、需要唱歌功能 → mimo-v2.5-tts(预置音色)
  • 需要独特音色 → mimo-v2.5-tts-voicedesign(文本描述生成)
  • 需要模仿特定声音 → mimo-v2.5-tts-voiceclone(音频样本复刻)

> 注意: TTS 有随机性,同样输入的效果可能不同,用户有需要时可以多生成几次以供挑选。

环境依赖

环境变量 说明 必需
MIMO_API_KEY MiMo API 密钥(MiMo 开放平台获取)

| 依赖 | 说明 | 必需

(truncated - see the full file via the links below)

Read as markdown · JSON record · Browse the source repository

File tree — 5 files
skills/mimo-v2-5-tts/SKILL.md
skills/mimo-v2-5-tts/scripts/feishu_send_audio.sh
skills/mimo-v2-5-tts/scripts/mimo_tts.py
skills/mimo-v2-5-tts/scripts/mimo_tts_voiceclone.py
skills/mimo-v2-5-tts/scripts/mimo_tts_voicedesign.py

Related skills

Tags

voice-cloning emotional-synthesis multilingual-tts voice-design character-voices dialect-support singing-synthesis director-mode audio-tagging preset-voices