$npx skillfedfor your agent

webrtcvad

Python interface to the Google WebRTC Voice Activity Detector (VAD)

With conditionsPyPI Artificial IntelligenceReleased Jan 2017480.4K downloads / moMITSource build

Decision gist · record as of 2026-08-14

sdist only — webrtcvad-2.0.10.tar.gz · builds from source
v2.0.10 · released 2017-01-07

Yes, with conditions. The underlying WebRTC VAD algorithm is mature and well-regarded. However, high install friction (compiled extension), dormancy since 2017-01-07, and uncertainty about modern Python compatibility mean you should verify it builds on your target platform and consider whether a more recently maintained alternative better suits your needs before committing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires 16-bit mono PCM audio at 8000, 16000, or 32000 Hz; frame duration must be 10, 20, or 30 ms.
  • Compilation of C extension may require build tools and development headers.
  • High install friction due to compiled C extension dependency (webrtcvad-2.0.10.tar.gz).

License · maintenance · safety

MIT (permissive) — MIT license (permissive) imposes no significant restrictions on use, modification, or distribution in commercial or private projects.

last release 2017-01-07 (3506 days) · last repo commit 2024-07-04 · 2,497 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 480,417 downloads/mo, #6,430 on PyPI

Verify before relying

pip install webrtcvad

import webrtcvad
vad = webrtcvad.Vad(1)
sample_rate = 16000
frame = b'\x00\x00' * (sample_rate * 10 // 1000)
result = vad.is_speech(frame, sample_rate)
  • Whether the package builds successfully on modern Python versions despite classifiers only listing up to 3.5.
  • Whether pre-built wheels are available to avoid compilation friction on common platforms.
  • Current real-world accuracy and performance compared to more recently maintained VAD alternatives.
Same gist for agents: .md · .json

What it is and what it does

webrtcvad is a Python binding to Google's WebRTC Voice Activity Detector, a classifier that determines whether short audio frames contain speech or silence. It accepts 16-bit mono PCM audio at fixed sample rates (8000, 16000, or 32000 Hz) in frames of 10, 20, or 30 milliseconds, and returns a boolean indicating whether speech is present. The detector supports aggressiveness levels (0–3) to tune sensitivity.

The package is used in speech recognition pipelines, telephony systems, and audio preprocessing workflows where you need to filter out silence or identify speech segments before further processing. It wraps a mature, well-regarded algorithm from the WebRTC project, but the Python wrapper itself has not been updated since 2017-01-07, creating uncertainty about compatibility with modern Python toolchains and whether better-maintained alternatives now exist.

Use it for

  • Preprocessing audio streams for automatic speech recognition by filtering out silence before sending to a speech-to-text service.
  • Segmenting recorded phone calls or voice messages to extract only the voiced portions for analysis or transcription.
  • Real-time voice activity detection in VoIP or conferencing applications to trigger recording or transmission only when speech is detected.
  • Training or evaluating speech detection models by labeling audio data as voiced or unvoiced at the frame level.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

The underlying WebRTC VAD algorithm is mature and well-regarded. However, high install friction (compiled extension), dormancy since 2017-01-07, and uncertainty about modern Python compatibility mean you should verify it builds on your target platform and consider whether a more recently maintained alternative better suits your needs before committing.

Install

webrtcvad on PyPI

Before you install

High install friction due to compiled C extension dependency (webrtcvad-2.0.10.tar.gz). The package is dormant—last release was 2017-01-07, over 3506 days ago—with no recent maintenance activity, though the repository remains active with 2497 stars.

Requires 16-bit mono PCM audio at 8000, 16000, or 32000 Hz; frame duration must be 10, 20, or 30 ms. Compilation of C extension may require build tools and development headers.

License in practice

MIT license (permissive) imposes no significant restrictions on use, modification, or distribution in commercial or private projects.

Quickstart

pip install webrtcvad

import webrtcvad
vad = webrtcvad.Vad(1)
sample_rate = 16000
frame = b'\x00\x00' * (sample_rate * 10 // 1000)
result = vad.is_speech(frame, sample_rate)

Verify before relying

  • Whether the package builds successfully on modern Python versions despite classifiers only listing up to 3.5.
  • Whether pre-built wheels are available to avoid compilation friction on common platforms.
  • Current real-world accuracy and performance compared to more recently maintained VAD alternatives.

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceDormant 3,506 days since the last release
Last repo commit
First released
Downloads480,417 / month, #6,430 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.2Programming Language :: Python :: 3.3Programming Language :: Python :: 3.4Programming Language :: Python :: 3.5Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Human Machine InterfacesTopic :: Scientific/Engineering :: Information Analysis

Evidence: webrtcvad-2.0.10.tar.gz

Tags

Capabilities
voice activity detectionspeech detection pythonwebrtc vadaudio classification voiced unvoicedspeech recognition preprocessingtelephony voice detection
Topics
audio-processingspeech-detection
PyPI keywords
speechrecognitionasrvoiceactivitydetectionvadwebrtc

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “speech detection python”

  • webrtcvadProvides a Python interface to Google's WebRTC Voice Activity…
  • silero-vadSilero VAD detects speech activity in audio files and streams,…
  • pymicro-vadA self-contained voice activity detector that processes 10ms chunks…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also pymicro-vad · silero-vad · webrtcvad-wheels · python_speech_features · realtimestt · aic-sdk · streamlit-webrtc · pyrnnoise · pytgcalls · voip-utils