$npx skillfedfor your agent

voxcpm

VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning

With conditionsPyPI Artificial IntelligenceReleased May 202699.7K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — voxcpm-2.0.3-py3-none-any.whl
v2.0.3 · released 2026-05-11 · Python >=3.10 · 23 runtime deps: torch, torchaudio, torchcodec, transformers, einops, gradio, inflect, addict

Yes, with conditions. VoxCPM2 is worth installing if you need multilingual TTS with voice cloning and design capabilities, have access to a modern GPU (CUDA ≥12.0), and can accommodate 23 dependencies including torch and transformers. The Apache-2.0 license permits commercial use. Active maintenance and no known vulnerabilities are positive signals. The main friction is the large dependency footprint and GPU requirement; it is not suitable for CPU-only or lightweight environments.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python ≥3.10 (<3.13), PyTorch ≥2.5.0, and CUDA ≥12.0; model weights must be downloaded from HuggingFace or ModelScope.
  • Low install friction with a pure-Python wheel, but carries 23 runtime dependencies including torch, transformers, and audio libraries.
  • Active maintenance with recent releases; requires Python ≥3.10 and PyTorch ≥2.5.0.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows free commercial use without restriction, making the package suitable for both open-source and proprietary applications.

last release 2026-05-11 (95 days) · last repo commit 2026-08-12 · 35,669 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 99,718 downloads/mo, #13,013 on PyPI

Verify before relying

pip install voxcpm

from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained("openbmb/VoxCPM2", load_denoiser=False)
wav = model.generate(text="Hello world", cfg_value=2.0, inference_timesteps=10)
sf.write("output.wav", wav, model.tts_model.sample_rate)
  • Actual inference speed (RTF) on consumer hardware; documentation cites RTF ~0.3 on RTX 4090 but real-world performance varies.
  • Memory requirements for loading the 2B parameter model on different GPU tiers.
  • Quality and naturalness of voice design outputs compared to reference-based cloning.
  • Multilingual synthesis quality parity across all 30 supported languages.
Same gist for agents: .md · .json

What it is and what it does

VoxCPM2 is a large language model–based text-to-speech system that bypasses traditional discrete tokenization, instead generating continuous speech representations end-to-end. It outputs 48kHz studio-quality audio and supports 30 languages including Chinese dialects, with no language tags required. The model is built on a MiniCPM-4 backbone and trained on over 2 million hours of multilingual speech data.

The package offers three primary synthesis modes: voice design (create a voice from natural-language description alone), controllable voice cloning (clone a voice from a short reference clip with optional style guidance), and ultimate cloning (provide both reference audio and transcript for seamless continuation with full vocal nuance preservation). It also supports streaming generation for real-time applications. All synthesis is context-aware, automatically inferring appropriate prosody and expressiveness from text content.

Use it for

  • Generate natural multilingual speech for applications serving users in 30 languages without language-specific models or tags.
  • Create synthetic voices from text descriptions for brand narration, character voices, or accessibility without reference audio.
  • Clone a speaker's voice from a short clip and adjust emotion, pace, or tone while preserving original timbre.
  • Build real-time speech synthesis pipelines using streaming generation with low latency on modern GPUs.
  • Reproduce exact vocal characteristics (timbre, rhythm, emotion) by providing reference audio and its transcript for high-fidelity cloning.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

VoxCPM2 is worth installing if you need multilingual TTS with voice cloning and design capabilities, have access to a modern GPU (CUDA ≥12.0), and can accommodate 23 dependencies including torch and transformers. The Apache-2.0 license permits commercial use. Active maintenance and no known vulnerabilities are positive signals. The main friction is the large dependency footprint and GPU requirement; it is not suitable for CPU-only or lightweight environments.

Install

voxcpm on PyPI

Before you install

Low install friction with a pure-Python wheel, but carries 23 runtime dependencies including torch, transformers, and audio libraries. Active maintenance with recent releases; requires Python ≥3.10 and PyTorch ≥2.5.0.

Requires Python ≥3.10 (<3.13), PyTorch ≥2.5.0, and CUDA ≥12.0; model weights must be downloaded from HuggingFace or ModelScope.

License in practice

Apache-2.0 permissive license allows free commercial use without restriction, making the package suitable for both open-source and proprietary applications.

Quickstart

pip install voxcpm

from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained("openbmb/VoxCPM2", load_denoiser=False)
wav = model.generate(text="Hello world", cfg_value=2.0, inference_timesteps=10)
sf.write("output.wav", wav, model.tts_model.sample_rate)

Verify before relying

  • Actual inference speed (RTF) on consumer hardware; documentation cites RTF ~0.3 on RTX 4090 but real-world performance varies.
  • Memory requirements for loading the 2B parameter model on different GPU tiers.
  • Quality and naturalness of voice design outputs compared to reference-based cloning.
  • Multilingual synthesis quality parity across all 30 supported languages.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
23 packages
torchtorchaudiotorchcodectransformerseinopsgradioinflectaddictwetextmodelscopedatasetshuggingface-hubpydantictqdmsimplejsonsortedcontainerssoundfilelibrosamatplotlibfunasrspacesargbindsafetensors
MaintenanceActively maintained 95 days since the last release
Last repo commit
First released
Downloads99,718 / month, #13,013 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11

Evidence: voxcpm-2.0.3-py3-none-any.whl

Tags

Capabilities
multilingual text to speechvoice cloning synthesisneural tts generationspeech synthesis with voice designhigh-quality audio generationcontrollable voice synthesisdiffusion-based tts
Topics
speech-synthesisvoice-cloningmultilingual-tts
PyPI keywords
voxcpmtext-to-speechttsspeech-synthesisvoice-cloningaideep-learningpytorch

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “voice cloning synthesis”

  • voxcpmVoxCPM2 is a tokenizer-free text-to-speech system that generates…
  • omnivoiceOmniVoice generates speech from text in over 600 languages using a…
  • TTSTTS is a deep learning library for text-to-speech synthesis that…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also chatterbox-tts · omnivoice · qwen-tts · pyworld · pocket-tts · TTS · coqui-tts · captcha · s3tokenizer · openai-whisper

Further reading