{"categories":[{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/7"}],"enrichment":{"capability":"Chatterbox TTS converts text to speech using open-source neural models, supporting English and 23+ languages with optional voice cloning from reference audio clips.","skillfed_tags":["speech-synthesis","voice-cloning","multilingual"],"use_cases":["Build low-latency voice agents or conversational AI that respond with natural speech in real time.","Generate multilingual narration for video, podcasts, or interactive media in 23+ languages with consistent voice.","Clone a specific speaker's voice from a short reference clip for personalized TTS without retraining.","Add expressive speech effects (laughter, coughing) to game dialogue, audiobooks, or creative projects.","Prototype TTS features before committing to a commercial service, using the open-source models locally."],"what_it_does":"Chatterbox TTS is a family of neural text-to-speech models from Resemble AI that convert written text into natural-sounding speech. The package includes three model variants: Turbo (350M parameters, English-only, optimized for low-latency voice agents), Multilingual (500M parameters, supports 23+ languages), and the original Chatterbox (500M parameters, English with creative control tuning). All models support zero-shot voice cloning\u2014you provide a reference audio clip and the model adapts its output to match that speaker's voice characteristics.\n\nThe package depends on a substantial ML stack: PyTorch, librosa for audio processing, transformers for language understanding, diffusers for generative modeling, and several specialized libraries (conformer, spacy-pkuseg for CJK text, pykakasi for Japanese). Every generated audio file includes an imperceptible neural watermark (Perth) for responsible AI tracking. The Turbo variant adds native support for paralinguistic tags like [cough] and [laugh] to inject realism, and reduces mel-spectrogram generation from 10 steps to one, trading some flexibility for speed. Configuration options (cfg_weight, exaggeration) allow tuning expressiveness and pacing.","worth_installing":"Yes, if you need open-source neural TTS with voice cloning and have GPU resources available. The package is actively maintained, permissively licensed, and offers competitive model variants. Install with caution on resource-constrained systems\u2014the dependency stack is heavy. For production voice agents requiring sub-200ms latency at scale, Resemble AI's commercial service is recommended as an alternative."},"id":"chatterbox-tts","links":{"html":"https://skillfed.io/packages/chatterbox-tts","md":"https://skillfed.io/packages/chatterbox-tts.md","pypi":"https://pypi.org/project/chatterbox-tts/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-03-26","license_spdx":null,"license_treatment":"permissive","name":"chatterbox-tts","python_support":"supports_current","summary":"Chatterbox: Open Source TTS and Voice Conversion by Resemble AI"},"popularity":{"monthly_downloads":203312,"position":9632,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.1.7"}
