{"enrichment":{"faq":[{"a":"Whisper is OpenAI's multilingual automatic speech recognition model trained on 680,000 hours of audio data. It transcribes speech to text across 99 languages, identifies the spoken language automatically, and can translate audio content to English. Whisper offers six configurable model sizes ranging from 39M to 1550M parameters, letting you balance accuracy against computational cost.","q":"What is Whisper speech to text and how does it work?"},{"a":"Yes. Whisper supports transcription and translation across 99 languages. You can transcribe audio in any of those languages to text in the original language, or translate the audio content to English. The model automatically detects which language is being spoken, making it ideal for multilingual podcasts, meetings, and international content.","q":"Can Whisper convert audio to text in multiple languages?"},{"a":"Whisper processes podcasts, meetings, and video audio into searchable transcripts with optional timestamped segments. Its robust handling of noisy audio, background noise, and varied speaking styles makes it well-suited for real-world recordings. You can run Whisper locally on your hardware or integrate it into transcription workflows without relying on cloud APIs.","q":"How can I use Whisper for podcast transcription and meeting notes?"},{"a":"Whisper is released under the Apache-2.0 license and available as open-source software. You can run it locally on your own hardware without cloud API costs or dependencies on external services. The model supports GPU acceleration for faster processing, and its six size options let you choose the right trade-off between speed and accuracy for your setup.","q":"Is Whisper open source and can I run it locally?"},{"a":"Whisper handles common audio formats including MP3, WAV, M4A, FLAC, and others. It can process audio from files, video content, or streaming sources. The model is designed to handle real-world audio quality, including background noise and poor recording conditions, making it practical for transcribing podcasts, meetings, and archival material.","q":"What audio formats and file types does Whisper support?"},{"a":"Whisper offers six model sizes from 39M to 1550M parameters. Smaller models run faster and use less memory, making them suitable for real-time or resource-constrained environments. Larger models deliver higher accuracy, especially on challenging audio. Choose based on your hardware, latency requirements, and accuracy needs\u2014GPU acceleration can speed up processing across all sizes.","q":"How do Whisper's different model sizes affect performance?"}],"shadow_tags":["audio-to-text","language-agnostic","offline-capable","open-source-asr","timestamp-generation","multilang-support","model-scaling","batch-processing","gpu-optimized","subtitle-generation"],"summary_rewrite":"Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support."},"files":[{"bytes":8106,"path":"skills/mlops/models/whisper/SKILL.md","sha256":"e81a517aa06f488c2216f1f0975104a57552c84a6a0d3a3ca69cd147b9bbde96","url":"https://skillfed.io/files/graniet/kheish/whisper/07f2a765/SKILL.md"}],"id":"graniet/kheish/whisper","links":{"html":"https://skillfed.io/graniet/kheish/whisper","md":"https://skillfed.io/graniet/kheish/whisper.md","repo":"https://github.com/graniet/kheish"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":22,"language":"Rust","last_updated":"2026-07-26","license":"Apache-2.0","name":"Whisper","publisher":"graniet","stars":264},"relations":{"similar":[{"id":"synthetic-sciences/openscience/whisper"},{"id":"Orchestra-Research/AI-Research-SKILLs/whisper"},{"id":"OpenLAIR/dr-claw/whisper"},{"id":"NousResearch/hermes-agent/whisper"},{"id":"ThePlasmak/faster-whisper/faster-whisper"},{"id":"chubbyguan/chubbyskills/podcast-transcribe"},{"id":"jamditis/claude-skills-journalism/video-transcribe"},{"id":"eyadsibai/ltk/multimodal-models"},{"id":"different-ai/agent-bank/video-subtitle-cutter"},{"id":"martinholovsky/claude-skills-generator/speech-to-text"}]},"slug":{"owner":"graniet","repo":"kheish","skill":"whisper"},"version":"07f2a765"}
