audiomentations
A Python library for audio data augmentation. Inspired by albumentations. Useful for machine learning.
Decision gist · record as of 2026-08-14
Yes. Audiomentations is actively maintained, has no security vulnerabilities, installs with low friction, and provides a well-documented, production-ready API for audio augmentation. It is particularly valuable if you are training audio ML models and need to generate synthetic variations without writing custom signal processing code. The MIT license removes legal barriers. Install it if audio data augmentation is part of your training pipeline.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction: pure Python wheel with no compiled dependencies beyond numpy and scipy.
- Actively maintained with a recent release (335 days ago) and 2310 repository stars.
- Supports Python 3.10–3.13.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted commercial and private use, modification, and distribution with minimal legal friction.
last release 2025-09-13 (335 days) · last repo commit 2026-04-13 · 2,310 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 211,335 downloads/mo, #9,481 on PyPI
Alternatives
Verify before relying
pip install audiomentations
from audiomentations import Compose, AddGaussianNoise, TimeStretch, PitchShift
import numpy as np
augment = Compose([
AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),
TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),
PitchShift(min_semitones=-4, max_semitones=4, p=0.5),
])
samples = np.random.uniform(low=-0.2, high=0.2, size=(32000,)).astype(np.float32)
augmented = augment(samples=samples, sample_rate=16000)- Whether all 7 runtime dependencies (numpy-minmax, numpy-rms, python-stretch, soxr) are required for core functionality or only for specific transforms.
- Performance characteristics and typical augmentation speed on CPU for real-time training pipelines.
- Multichannel audio support details and any channel-count constraints.
What it is and what it does
Audiomentations is a Python library for augmenting audio data by applying randomized transformations to waveforms. It provides a Compose-based API inspired by albumentations, letting you chain transforms like noise injection, pitch shifting, time stretching, filtering, and distortion with configurable probability. The library runs on CPU, supports both mono and multichannel audio, and integrates directly into training loops for TensorFlow/Keras and PyTorch.
The package solves the problem of generating synthetic audio variations during training to improve model robustness and generalization. Rather than manually coding audio effects, you declare a pipeline of transforms with parameter ranges, and the library applies them randomly to each sample during training. It is designed for practitioners building real-world audio ML systems—speech recognition, audio classification, sound event detection—where models trained only on clean lab data often fail in production.
Use it for
- Augment speech data during training to improve robustness to background noise and acoustic variation.
- Generate synthetic pitch and tempo variations to train music information retrieval models.
- Simulate room acoustics and microphone artifacts to make audio classifiers generalize across recording conditions.
- Add realistic distortion and compression artifacts to train models that handle low-quality or compressed audio.
- Combine multiple transforms in a pipeline to create diverse training samples from a small audio dataset.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Audiomentations is actively maintained, has no security vulnerabilities, installs with low friction, and provides a well-documented, production-ready API for audio augmentation. It is particularly valuable if you are training audio ML models and need to generate synthetic variations without writing custom signal processing code. The MIT license removes legal barriers. Install it if audio data augmentation is part of your training pipeline.
Install
audiomentations on PyPI
Before you install
Low friction: pure Python wheel with no compiled dependencies beyond numpy and scipy. Actively maintained with a recent release (335 days ago) and 2310 repository stars. Supports Python 3.10–3.13.
License in practice
MIT license permits unrestricted commercial and private use, modification, and distribution with minimal legal friction.
Quickstart
pip install audiomentations
from audiomentations import Compose, AddGaussianNoise, TimeStretch, PitchShift
import numpy as np
augment = Compose([
AddGaussianNoise(min_amplitude=0.001, max_amplitude=0.015, p=0.5),
TimeStretch(min_rate=0.8, max_rate=1.25, p=0.5),
PitchShift(min_semitones=-4, max_semitones=4, p=0.5),
])
samples = np.random.uniform(low=-0.2, high=0.2, size=(32000,)).astype(np.float32)
augmented = augment(samples=samples, sample_rate=16000)
Verify before relying
- Whether all 7 runtime dependencies (numpy-minmax, numpy-rms, python-stretch, soxr) are required for core functionality or only for specific transforms.
- Performance characteristics and typical augmentation speed on CPU for real-time training pipelines.
- Multichannel audio support details and any channel-count constraints.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagesnumpynumpy-minmaxnumpy-rmslibrosapython-stretchscipysoxr |
| Maintenance | Actively maintained 335 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 211,335 / month, #9,481 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: MultimediaTopic :: Multimedia :: Sound/AudioTopic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: audiomentations-0.43.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “audio transformations for machine learning”
- audiomentationsAudiomentations applies randomized audio transformations—noise…
- nlpaugnlpaug generates synthetic augmented text and audio data for machine…
- audiolabaudiolab loads, processes, and saves audio files in multiple formats…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also nlpaug · pedalboard · torch-audiomentations · descript-audiotools · noisereduce · python-stretch · stftpitchshift · sox · albumentations · imgaug