panns-inference
panns_inference: audio tagging and sound event detection inference toolbox
Decision gist · record as of 2026-08-14
Yes, if you need audio tagging or sound event detection and can tolerate dormant maintenance. The package is stable, has no known vulnerabilities, and low install friction. However, do not expect bug fixes or updates—verify that the pretrained models and PyTorch compatibility meet your production requirements before committing to it for critical systems.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- PyTorch >= 1.0 is required; GPU device support (cuda) is optional but models run on CPU if device='cpu' is specified.
- Low install friction with a pure-Python wheel.
- Maintenance is dormant—last commit was 2024-03-05 and no releases since 2023-03-26—so expect no active bug fixes or updates, though the codebase remains archived and available.
License · maintenance · safety
permissive license (permissive) — MIT license (permissive) means you can use, modify, and distribute this package freely with minimal restrictions, provided you retain the license notice.
last release 2023-03-26 (1237 days) · last repo commit 2024-03-05 · 268 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 81,794 downloads/mo, #14,200 on PyPI
Alternatives
Verify before relying
import librosa
from panns_inference import AudioTagging, SoundEventDetection
audio_path = 'examples/audio.wav'
audio, _ = librosa.core.load(audio_path, sr=32000, mono=True)
audio = audio[None, :]
at = AudioTagging(checkpoint_path=None, device='cuda')
clipwise_output, embedding = at.inference(audio)- Whether pretrained checkpoint files are automatically downloaded or must be manually provided.
- Current accuracy/performance metrics for the bundled models on modern audio datasets.
- Compatibility with recent PyTorch and librosa versions beyond the minimum requirements.
What it is and what it does
panns_inference wraps pretrained audio neural networks from the PANNs project to perform two main tasks: audio tagging (labeling what sounds are present in an audio clip with confidence scores) and sound event detection (identifying when specific sounds occur within an audio file). It loads audio via librosa, feeds it through pretrained convolutional neural network models, and returns either clip-level tags with probabilities or frame-level event predictions. The package is designed for researchers and developers who want to apply pretrained audio understanding without training their own models.
The package depends on librosa for audio loading and preprocessing, matplotlib for visualization, and torchlibrosa for PyTorch-compatible audio feature extraction. It requires PyTorch >= 1.0 and Python >= 3.6. Since the last release was in March 2023 and the repository shows no recent commits, the package is in maintenance limbo—it will work with existing code but should not be expected to receive updates for compatibility with newer dependencies.
Use it for
- Classify audio clips into semantic categories (speech, music, vehicle sounds, etc.) with confidence scores for content moderation or audio organization.
- Detect and timestamp specific sound events within longer audio recordings for audio annotation or event-based analysis.
- Extract audio embeddings from pretrained models for downstream machine learning tasks like clustering or similarity search.
- Prototype audio understanding features in applications without the overhead of training custom models from scratch.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need audio tagging or sound event detection and can tolerate dormant maintenance.
The package is stable, has no known vulnerabilities, and low install friction. However, do not expect bug fixes or updates—verify that the pretrained models and PyTorch compatibility meet your production requirements before committing to it for critical systems.
Install
panns-inference on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance is dormant—last commit was 2024-03-05 and no releases since 2023-03-26—so expect no active bug fixes or updates, though the codebase remains archived and available.
PyTorch >= 1.0 is required; GPU device support (cuda) is optional but models run on CPU if device='cpu' is specified.
License in practice
MIT license (permissive) means you can use, modify, and distribute this package freely with minimal restrictions, provided you retain the license notice.
Quickstart
import librosa
from panns_inference import AudioTagging, SoundEventDetection
audio_path = 'examples/audio.wav'
audio, _ = librosa.core.load(audio_path, sr=32000, mono=True)
audio = audio[None, :]
at = AudioTagging(checkpoint_path=None, device='cuda')
clipwise_output, embedding = at.inference(audio)
Verify before relying
- Whether pretrained checkpoint files are automatically downloaded or must be manually provided.
- Current accuracy/performance metrics for the bundled models on modern audio datasets.
- Compatibility with recent PyTorch and librosa versions beyond the minimum requirements.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.6 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesmatplotliblibrosatorchlibrosa |
| Maintenance | Dormant 1,237 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 81,794 / month, #14,200 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: panns_inference-0.1.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “audio tagging inference”
- panns-inferenceProvides pretrained neural network models for audio tagging and sound…
- mediafileMediaFile reads and writes metadata tags across many audio file…
- funasrFunASR is a speech recognition toolkit that transcribes audio offline…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also resemble-perth · snac · laion-clap · speechbrain · sherpa-onnx-core · aubio · encodec · torchcrepe · pyacoustid · audio-separator