webrtcvad-wheels
Python interface to the Google WebRTC Voice Activity Detector (VAD) [released with binary wheels!]
Decision gist · record as of 2026-08-14
Yes. The package is stable (Production/Stable status), has no known vulnerabilities, zero runtime dependencies, and ships with pre-built wheels that eliminate compilation friction on major platforms. The aging maintenance status is acceptable for a mature, focused tool with a narrow scope. Install it if you need reliable voice activity detection in Python.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Audio must be 16-bit mono PCM at 8000, 16000, 32000, or 48000 Hz; frames must be exactly 10, 20, or 30 ms in duration.
- Medium install friction due to compiled C extension, but mitigated by pre-built wheels covering Windows, macOS, and Linux across multiple architectures.
- Repository is aging (708 days since last release) but remains actively maintained with recent commits.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you must include the license text in distributions.
last release 2024-09-05 (708 days) · last repo commit 2026-01-12 · 41 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 528,543 downloads/mo, #6,164 on PyPI
Alternatives
Verify before relying
pip install webrtcvad-wheels
import webrtcvad
vad = webrtcvad.Vad(1) # aggressiveness 0-3
frame = b'\x00\x00' * int(16000 * 10 / 1000) # 10ms at 16kHz
is_speech = vad.is_speech(frame, 16000)- Whether the package supports Python versions below 3.6 despite classifiers listing 3.6+
- Exact memory leak fixes and performance improvements claimed in version history beyond what the fact sheet documents
- Specific audio frame duration requirements and sample rate constraints beyond the documented 10, 20, or 30 ms frames
What it is and what it does
This package wraps Google's WebRTC Voice Activity Detector, a fast and accurate classifier that determines whether a short audio segment contains speech or silence. It accepts 16-bit mono PCM audio at standard sample rates and returns a boolean result, with tunable aggressiveness (0–3) to control sensitivity to non-speech sounds. The package is distributed with pre-compiled binary wheels, eliminating the need to build the underlying C extension on Windows, macOS, and Linux across multiple architectures and CPU types.
The VAD is commonly used as a preprocessing step in speech recognition pipelines, telephony systems, and audio analysis workflows where you need to filter out silence or identify voiced regions before further processing. With no runtime dependencies and a simple API, it integrates easily into Python audio applications.
Use it for
- Filter silence from audio recordings before sending to a speech-to-text service to reduce processing cost and latency.
- Segment a continuous audio stream into voiced and unvoiced regions for telephony or voice call analysis.
- Preprocess microphone input in real-time speech recognition to skip processing during silence.
- Detect speech activity in surveillance or meeting recordings to identify when participants are speaking.
- Build a voice activity detector for audio quality assessment or speaker diarization pipelines.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is stable (Production/Stable status), has no known vulnerabilities, zero runtime dependencies, and ships with pre-built wheels that eliminate compilation friction on major platforms. The aging maintenance status is acceptable for a mature, focused tool with a narrow scope. Install it if you need reliable voice activity detection in Python.
Install
webrtcvad-wheels on PyPI
Before you install
Medium install friction due to compiled C extension, but mitigated by pre-built wheels covering Windows, macOS, and Linux across multiple architectures. Repository is aging (708 days since last release) but remains actively maintained with recent commits.
Audio must be 16-bit mono PCM at 8000, 16000, 32000, or 48000 Hz; frames must be exactly 10, 20, or 30 ms in duration.
License in practice
MIT license permits commercial and private use with minimal restrictions; you must include the license text in distributions.
Quickstart
pip install webrtcvad-wheels
import webrtcvad
vad = webrtcvad.Vad(1) # aggressiveness 0-3
frame = b'\x00\x00' * int(16000 * 10 / 1000) # 10ms at 16kHz
is_speech = vad.is_speech(frame, 16000)
Verify before relying
- Whether the package supports Python versions below 3.6 despite classifiers listing 3.6+
- Exact memory leak fixes and performance improvements claimed in version history beyond what the fact sheet documents
- Specific audio frame duration requirements and sample rate constraints beyond the documented 10, 20, or 30 ms frames
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Aging 708 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 528,543 / month, #6,164 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Human Machine InterfacesTopic :: Scientific/Engineering :: Information Analysis |
Evidence: webrtcvad_wheels-2.0.14-cp310-cp310-macosx_10_9_x86_64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-macosx_11_0_arm64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl; webrtcvad_wheels-2.0.14-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-manylinux_2_5_i686.manylinux1_i686.manylinux_2_17_i686.manylinux2014_i686.whl; webrtcvad_wheels-2.0.14-cp310-cp310-musllinux_1_2_aarch64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-musllinux_1_2_i686.whl; webrtcvad_wheels-2.0.14-cp310-cp310-musllinux_1_2_ppc64le.whl; webrtcvad_wheels-2.0.14-cp310-cp310-musllinux_1_2_x86_64.whl; webrtcvad_wheels-2.0.14-cp310-cp310-win32.whl; webrtcvad_wheels-2.0.14-cp310-cp310-win_amd64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-macosx_10_9_x86_64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-macosx_11_0_arm64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl; webrtcvad_wheels-2.0.14-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-manylinux_2_5_i686.manylinux1_i686.manylinux_2_17_i686.manylinux2014_i686.whl; webrtcvad_wheels-2.0.14-cp311-cp311-musllinux_1_2_aarch64.whl; webrtcvad_wheels-2.0.14-cp311-cp311-musllinux_1_2_i686.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “vad webrtc python”
- webrtcvad-wheelsDetects voiced versus unvoiced segments in audio by wrapping Google's…
- webrtcvadProvides a Python interface to Google's WebRTC Voice Activity…
- livekit-plugins-assemblyaiIntegrates AssemblyAI speech-to-text into the LiveKit Agents…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also webrtcvad · pymicro-vad · silero-vad · livekit-plugins-turn-detector · realtimestt · pyrnnoise · aic-sdk · openwakeword · whisper-timestamped · pipecat-ai-small-webrtc-prebuilt