coqui-tts
Deep learning for Text to Speech.
Decision gist · record as of 2026-08-14
Yes, if you need multilingual text-to-speech synthesis or voice conversion. The library is actively maintained, has no known vulnerabilities, and offers both easy inference and advanced training capabilities. Install it if you want pretrained models out-of-the-box or plan to fine-tune. The MPL-2.0 copyleft license is permissive for unmodified use in closed-source work but requires sharing modifications. Be aware that external PyTorch installation is required and the dependency footprint is large.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10–3.14.
- External PyTorch installation required (not bundled since 0.27.4).
- GPU support optional but recommended for inference speed.
License · maintenance · safety
MPL-2.0 (copyleft) — MPL-2.0 is copyleft: derivative works and modifications must be distributed under the same license. If you modify the library itself, you must share those changes. Using it unmodified in closed-source applications is permitted.
last release 2026-01-26 (200 days) · last repo commit 2026-06-10 · 2,312 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 183,668 downloads/mo, #10,063 on PyPI
Alternatives
Verify before relying
pip install coqui-tts
from TTS.api import TTS
tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC", gpu=False)
tts.tts_to_file(text="Hello world", file_path="output.wav")- Whether the 1100 languages claim refers to Fairseq models or all available models in the library
- Inference latency and memory requirements for typical use cases
- Whether voice cloning (mentioned in 0.27.0 news) requires additional setup or training data
- Exact PyTorch version requirements and compatibility with PyTorch 2.2+
What it is and what it does
Coqui TTS is a deep learning library for converting text into natural-sounding speech. It bundles multiple neural architectures (Tacotron2, Glow-TTS, VITS, XTTS, and others) with pretrained weights across many languages, plus vocoders to convert spectrograms to audio. You can use it off-the-shelf for inference, fine-tune existing models on your own data, or train new models from scratch. It also supports voice conversion (changing a speaker's identity while preserving content) and voice cloning with minimal reference audio.
The library is designed for both research and production use. It provides command-line tools and a Python API, with utilities for dataset curation and analysis. The main constraint is that external PyTorch installation is required, and the full dependency stack (transformers, librosa, scipy, numba, and others) is substantial. Training new models is computationally expensive; inference can run on CPU but is much faster on GPU.
Use it for
- Generate speech from text in multiple languages using pretrained models without training
- Fine-tune an existing TTS model on your own voice or dataset to customize output
- Build a voice cloning system that generates speech in a target speaker's voice from a short audio sample
- Convert one speaker's voice to another while preserving the linguistic content
- Train a multilingual or multi-speaker TTS model from scratch on custom data
- Analyze and curate TTS training datasets using built-in dataset analysis tools
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need multilingual text-to-speech synthesis or voice conversion.
The library is actively maintained, has no known vulnerabilities, and offers both easy inference and advanced training capabilities. Install it if you want pretrained models out-of-the-box or plan to fine-tune. The MPL-2.0 copyleft license is permissive for unmodified use in closed-source work but requires sharing modifications. Be aware that external PyTorch installation is required and the dependency footprint is large.
Install
coqui-tts on PyPI
Before you install
Low friction installation via wheel distribution. Active maintenance with last commit 2026-06-10. Requires Python 3.10–3.14. The 21 runtime dependencies include heavy scientific stacks (transformers, librosa, scipy, numba) typical of ML libraries.
Requires Python 3.10–3.14. External PyTorch installation required (not bundled since 0.27.4). GPU support optional but recommended for inference speed.
License in practice
MPL-2.0 is copyleft: derivative works and modifications must be distributed under the same license. If you modify the library itself, you must share those changes. Using it unmodified in closed-source applications is permitted.
Quickstart
pip install coqui-tts
from TTS.api import TTS
tts = TTS(model_name="tts_models/en/ljspeech/tacotron2-DDC", gpu=False)
tts.tts_to_file(text="Hello world", file_path="output.wav")
Verify before relying
- Whether the 1100 languages claim refers to Fairseq models or all available models in the library
- Inference latency and memory requirements for typical use cases
- Whether voice cloning (mentioned in 0.27.0 news) requires additional setup or training data
- Exact PyTorch version requirements and compatibility with PyTorch 2.2+
Package facts
| License | MPL-2.0 copyleft |
| Python support | Supports the current Python release <3.15,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 21 packagesanyasciicoqpit-configcoqui-tts-trainereinopsfsspecinflectko-speech-toolslibrosamatplotlibmonotonic-alignment-searchnum2wordsnumbanumpypackagingpysbdpyyamlscipysoundfiletqdmtransformerstyping-extensions |
| Maintenance | Actively maintained 200 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 183,668 / month, #10,063 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Mozilla Public License 2.0 (MPL 2.0)Operating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: MultimediaTopic :: Multimedia :: Sound/AudioTopic :: Multimedia :: Sound/Audio :: SpeechTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software DevelopmentTopic :: Software Development :: Libraries :: Python Modules |
Evidence: coqui_tts-0.27.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “voice generation deep learning”
- coqui-ttsCoqui TTS synthesizes speech from text using deep learning models,…
- TTSTTS is a deep learning library for text-to-speech synthesis that…
- ResemblyzerResemblyzer generates a 256-value embedding that summarizes voice…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also TTS · chatterbox-tts · f5-tts · kokoro-onnx · omnivoice · piper-tts · silero · kokoro · speechbrain · pocket-tts