{"categories":[{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/10"}],"enrichment":{"capability":"VoxCPM2 is a tokenizer-free text-to-speech system that generates multilingual speech directly via diffusion-autoregressive architecture, supporting voice design, controllable voice cloning, and 48kHz audio output across 30 languages.","skillfed_tags":["speech-synthesis","voice-cloning","multilingual-tts"],"use_cases":["Generate natural multilingual speech for applications serving users in 30 languages without language-specific models or tags.","Create synthetic voices from text descriptions for brand narration, character voices, or accessibility without reference audio.","Clone a speaker's voice from a short clip and adjust emotion, pace, or tone while preserving original timbre.","Build real-time speech synthesis pipelines using streaming generation with low latency on modern GPUs.","Reproduce exact vocal characteristics (timbre, rhythm, emotion) by providing reference audio and its transcript for high-fidelity cloning."],"what_it_does":"VoxCPM2 is a large language model\u2013based text-to-speech system that bypasses traditional discrete tokenization, instead generating continuous speech representations end-to-end. It outputs 48kHz studio-quality audio and supports 30 languages including Chinese dialects, with no language tags required. The model is built on a MiniCPM-4 backbone and trained on over 2 million hours of multilingual speech data.\n\nThe package offers three primary synthesis modes: voice design (create a voice from natural-language description alone), controllable voice cloning (clone a voice from a short reference clip with optional style guidance), and ultimate cloning (provide both reference audio and transcript for seamless continuation with full vocal nuance preservation). It also supports streaming generation for real-time applications. All synthesis is context-aware, automatically inferring appropriate prosody and expressiveness from text content.","worth_installing":"Yes, with conditions. VoxCPM2 is worth installing if you need multilingual TTS with voice cloning and design capabilities, have access to a modern GPU (CUDA \u226512.0), and can accommodate 23 dependencies including torch and transformers. The Apache-2.0 license permits commercial use. Active maintenance and no known vulnerabilities are positive signals. The main friction is the large dependency footprint and GPU requirement; it is not suitable for CPU-only or lightweight environments."},"id":"voxcpm","links":{"html":"https://skillfed.io/packages/voxcpm","md":"https://skillfed.io/packages/voxcpm.md","pypi":"https://pypi.org/project/voxcpm/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-05-11","license_spdx":"Apache-2.0","license_treatment":"permissive","name":"voxcpm","python_support":"supports_current","summary":"VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning"},"popularity":{"monthly_downloads":99718,"position":13013,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"2.0.3"}
