audio-separator
Easy to use audio stem separation, using various models from UVR trained primarily by @Anjok07
Decision gist · record as of 2026-08-14
Yes, if you need audio stem separation and can accommodate the dependency footprint. The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers both CLI and library interfaces. Install friction is low despite 22 dependencies. The main gotcha is the separate FFmpeg system requirement and the need to choose an appropriate hardware acceleration path (CUDA, CoreML, CPU, or experimental DirectML). Recommended for music production, karaoke generation, and audio research workflows.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires FFmpeg to be installed separately on the system; torch and ONNX runtime must be available; GPU acceleration (CUDA 11.8 or 12.2, CoreML, or DirectML) is optional but recommended for performance.
- Low friction install with a pure-Python wheel.
- Requires torch and 22 runtime dependencies including audio libraries (librosa, soundfile, pydub) and ONNX tooling.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license but face no copyleft obligations.
last release 2026-07-20 (25 days) · last repo commit 2026-07-20 · 1,314 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 363,613 downloads/mo, #7,219 on PyPI
Alternatives
Verify before relying
pip install audio-separator
from audio_separator.separator import Separator
separator = Separator(model_name="bs_roformer")
separator.separate("input.wav")- Whether the package handles all common audio formats (WAV, MP3, FLAC, M4A) as claimed in the description.
- Performance characteristics and typical separation quality across different model architectures (MDX-Net, VR Arch, Demucs, MDXC).
- Memory requirements for inference on typical consumer hardware.
- Compatibility and stability of DirectML acceleration on AMD and Intel GPUs.
What it is and what it does
Audio Separator is a Python package that uses pre-trained deep learning models to decompose audio files into separate stems—such as vocals, drums, bass, and other instruments. It wraps models from Ultimate Vocal Remover (UVR) trained by @Anjok07, making them accessible both as a command-line tool and as a library for batch processing or integration into larger audio workflows. The package supports multiple model architectures (MDX-Net, VR Arch, Demucs, MDXC/RoFormer) and can run on CPU or accelerated hardware (NVIDIA CUDA, Apple Silicon CoreML, or experimental Windows DirectML).
The most common use case is separating a song into instrumental and vocal stems for karaoke production, but the underlying models support more granular decomposition into drums, bass, piano, guitar, and other sources. The package also handles audio preprocessing and format conversion via FFmpeg, making it straightforward to process diverse input files. With 22 runtime dependencies including torch, librosa, and ONNX tooling, it trades installation complexity for a complete, batteries-included audio processing pipeline.
Use it for
- Generate karaoke tracks by separating vocals from instrumental stems in music files.
- Extract individual instrument stems (drums, bass, guitar) from recordings for remixing or music production.
- Denoise or remove echo/reverb from audio recordings using specialized UVR models.
- Batch process large music libraries to create stem versions for archival or reuse.
- Integrate stem separation into a larger audio analysis or music information retrieval pipeline.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need audio stem separation and can accommodate the dependency footprint.
The package is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers both CLI and library interfaces. Install friction is low despite 22 dependencies. The main gotcha is the separate FFmpeg system requirement and the need to choose an appropriate hardware acceleration path (CUDA, CoreML, CPU, or experimental DirectML). Recommended for music production, karaoke generation, and audio research workflows.
Install
audio-separator on PyPI
Before you install
Low friction install with a pure-Python wheel. Requires torch and 22 runtime dependencies including audio libraries (librosa, soundfile, pydub) and ONNX tooling. Active maintenance with recent releases.
Requires FFmpeg to be installed separately on the system; torch and ONNX runtime must be available; GPU acceleration (CUDA 11.8 or 12.2, CoreML, or DirectML) is optional but recommended for performance.
License in practice
MIT license permits commercial and private use with minimal restrictions; you must include a copy of the license but face no copyleft obligations.
Quickstart
pip install audio-separator
from audio_separator.separator import Separator
separator = Separator(model_name="bs_roformer")
separator.separate("input.wav")
Verify before relying
- Whether the package handles all common audio formats (WAV, MP3, FLAC, M4A) as claimed in the description.
- Performance characteristics and typical separation quality across different model architectures (MDX-Net, VR Arch, Demucs, MDXC).
- Memory requirements for inference on typical consumer hardware.
- Compatibility and stability of DirectML acceleration on AMD and Intel GPUs.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 22 packagesaudioop-ltsbeartypediffqdiffq-fixedeinopsjuliuslibrosaml_collectionsnumpyonnx-weeklyonnx2torch-py313pydubpyyamlrequestsresampyrotary-embedding-torchsampleratescipysixsoundfiletorchtqdm |
| Maintenance | Actively maintained 25 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 363,613 / month, #7,219 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Multimedia :: Sound/AudioTopic :: Multimedia :: Sound/Audio :: Mixers |
Evidence: audio_separator-0.44.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “audio stem separation”
- audio-separatorSeparates audio files into multiple stems (vocals, instruments,…
- openunmixSeparates music into individual stems (vocals, drums, bass, other…
- demucsDemucs separates music into individual stems—drums, bass, vocals, and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Sound/Audio packages
Reads and writes audio files in formats like WAV, FLAC, OGG, and MAT through libsndfile, exposing audio data as NumPy arrays.
Pydub provides a high-level Python interface for loading, manipulating, and exporting audio files with simple operations like slicing, concatenation, and format conversion.
However, the abandoned status since 2021-03-10 means no future fixes or compatibility updates—use it only if you can tolerate potential issues with newer Python…
Mutagen reads and writes audio metadata (tags) across many audio formats including MP3, FLAC, OGG, MP4, WavPack, and others, supporting ID3v2 and APEv2 tag editing.
Install it if you need to read or edit audio tags programmatically.
Provides PyTorch-based audio processing, transforms, and dataloaders for machine learning tasks, with GPU acceleration and autograd support for trainable audio features.
MoviePy is a Python library for video editing that reads, processes, and writes video and audio files by converting them to numpy arrays for frame-level manipulation and effect application.
Reads metadata (artist, title, duration, bitrate, and more) from audio files in formats including MP3, MP4, FLAC, OGG, WAV, and others, without writing or modifying tags.
Install it if you need to extract metadata from audio files without the overhead of a heavier library.
See also demucs · openunmix · asteroid-filterbanks · audiofile · panns-inference · resemble-perth · snac · torchcodec · decord2