$npx skillfedfor your agent

chatterbox-tts

Chatterbox: Open Source TTS and Voice Conversion by Resemble AI

With conditionsPyPI Artificial IntelligenceReleased Mar 2026203.3K downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — chatterbox_tts-0.1.7-py3-none-any.whl
v0.1.7 · released 2026-03-26 · Python >=3.10 · 15 runtime deps: numpy, librosa, s3tokenizer, torch, torchaudio, transformers, diffusers, resemble-perth

Yes, if you need open-source neural TTS with voice cloning and have GPU resources available. The package is actively maintained, permissively licensed, and offers competitive model variants. Install with caution on resource-constrained systems—the dependency stack is heavy. For production voice agents requiring sub-200ms latency at scale, Resemble AI's commercial service is recommended as an alternative.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.10.
  • GPU (CUDA) strongly recommended for practical inference speed.
  • A reference audio clip (approximately 10 seconds) is needed for voice cloning.

License · maintenance · safety

permissive license (permissive) — MIT License permits free use, modification, and distribution with minimal restrictions. You may use this in commercial projects and modify the code as needed, provided you retain the copyright notice and license text.

last release 2026-03-26 (141 days) · last repo commit 2026-07-21 · 25,990 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 203,312 downloads/mo, #9,632 on PyPI

Verify before relying

pip install chatterbox-tts

import torchaudio as ta
from chatterbox.tts_turbo import ChatterboxTurboTTS

model = ChatterboxTurboTTS.from_pretrained(device="cuda")
wav = model.generate("Hello world", audio_prompt_path="ref.wav")
ta.save("output.wav", wav, model.sr)
  • Whether the Turbo model's single-step mel decoder actually maintains audio fidelity compared to multi-step alternatives in blind listening tests.
  • Actual inference latency and VRAM requirements for the Turbo model on consumer GPUs.
  • Whether paralinguistic tags ([cough], [laugh], etc.) work reliably across all supported languages or only English.
  • Compatibility and performance on non-CUDA devices (CPU, AMD ROCm, Apple Metal).
Same gist for agents: .md · .json

What it is and what it does

Chatterbox TTS is a family of neural text-to-speech models from Resemble AI that convert written text into natural-sounding speech. The package includes three model variants: Turbo (350M parameters, English-only, optimized for low-latency voice agents), Multilingual (500M parameters, supports 23+ languages), and the original Chatterbox (500M parameters, English with creative control tuning). All models support zero-shot voice cloning—you provide a reference audio clip and the model adapts its output to match that speaker's voice characteristics.

The package depends on a substantial ML stack: PyTorch, librosa for audio processing, transformers for language understanding, diffusers for generative modeling, and several specialized libraries (conformer, spacy-pkuseg for CJK text, pykakasi for Japanese). Every generated audio file includes an imperceptible neural watermark (Perth) for responsible AI tracking. The Turbo variant adds native support for paralinguistic tags like [cough] and [laugh] to inject realism, and reduces mel-spectrogram generation from 10 steps to one, trading some flexibility for speed. Configuration options (cfg_weight, exaggeration) allow tuning expressiveness and pacing.

Use it for

  • Build low-latency voice agents or conversational AI that respond with natural speech in real time.
  • Generate multilingual narration for video, podcasts, or interactive media in 23+ languages with consistent voice.
  • Clone a specific speaker's voice from a short reference clip for personalized TTS without retraining.
  • Add expressive speech effects (laughter, coughing) to game dialogue, audiobooks, or creative projects.
  • Prototype TTS features before committing to a commercial service, using the open-source models locally.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need open-source neural TTS with voice cloning and have GPU resources available.

The package is actively maintained, permissively licensed, and offers competitive model variants. Install with caution on resource-constrained systems—the dependency stack is heavy. For production voice agents requiring sub-200ms latency at scale, Resemble AI's commercial service is recommended as an alternative.

Install

chatterbox-tts on PyPI

Before you install

Low install friction with a pure Python wheel. The package is actively maintained with recent commits and substantial community interest (25990 GitHub stars). However, it carries 15 runtime dependencies including torch, torchaudio, transformers, and diffusers—a heavy ML stack that will require significant disk and memory resources during installation.

Requires Python >=3.10. GPU (CUDA) strongly recommended for practical inference speed. A reference audio clip (approximately 10 seconds) is needed for voice cloning. The full dependency stack (torch, transformers, diffusers) will consume several gigabytes of disk space and VRAM.

License in practice

MIT License permits free use, modification, and distribution with minimal restrictions. You may use this in commercial projects and modify the code as needed, provided you retain the copyright notice and license text.

Quickstart

pip install chatterbox-tts

import torchaudio as ta
from chatterbox.tts_turbo import ChatterboxTurboTTS

model = ChatterboxTurboTTS.from_pretrained(device="cuda")
wav = model.generate("Hello world", audio_prompt_path="ref.wav")
ta.save("output.wav", wav, model.sr)

Verify before relying

  • Whether the Turbo model's single-step mel decoder actually maintains audio fidelity compared to multi-step alternatives in blind listening tests.
  • Actual inference latency and VRAM requirements for the Turbo model on consumer GPUs.
  • Whether paralinguistic tags ([cough], [laugh], etc.) work reliably across all supported languages or only English.
  • Compatibility and performance on non-CUDA devices (CPU, AMD ROCm, Apple Metal).

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
15 packages
numpylibrosas3tokenizertorchtorchaudiotransformersdiffusersresemble-perthconformersafetensorsspacy-pkusegpykakasigradiopyloudnormomegaconf
MaintenanceActively maintained 141 days since the last release
Last repo commit
First released
Downloads203,312 / month, #9,632 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: chatterbox_tts-0.1.7-py3-none-any.whl

Tags

Capabilities
text to speech neural modelsmultilingual tts synthesisvoice cloning from audiozero-shot speech generationparalinguistic speech tagslow-latency tts inferenceopen source tts models
Topics
speech-synthesisvoice-cloningmultilingual

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “text to speech neural models”

  • chatterbox-ttsChatterbox TTS converts text to speech using open-source neural…
  • TTSTTS is a deep learning library for text-to-speech synthesis that…
  • sileroSilero provides pre-trained text-to-speech models that convert text…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also voxcpm · TTS · coqui-tts · omnivoice · mlx-audio · qwen-tts · gTTS · pocket-tts · pyworld · kokoro