$npx skillfedfor your agent

Speech To Text

This skill implements speech-to-text using Faster Whisper for converting audio input into written transcriptions. It prioritizes local processing, immediate deletion of audio data, and secure handling of voice information while supporting real-time streaming, multiple languages, and hardware-optimized model selection.

Speech To Text converts spoken audio into written text using Faster Whisper with privacy-first local processing.

AI-generated summary based on this skill's SKILL.md

45 4 Unlicenseupdated by martinholovsky

Decision gist · record as of 2025-12-06

Speech To Text converts spoken audio into written text using Faster Whisper with privacy-first local processing. This skill implements speech-to-text using Faster Whisper for converting audio input into written transcriptions. It prioritizes local processing, immediate deletion of audio data, and secure handling of voice information while supporting real-time streaming, multiple languages, and hardware-optimized model selection.

manual: git clone https://github.com/martinholovsky/claude-skills-generator → cp -r claude-skills-generator ~/.claude/skills/speech-to-text

Use it when

  • Speech To Text accepts audio input and transcribes it into written text.
  • Yes, Speech To Text transcribes speech for accessibility and documentation purposes.
Same gist for agents: .md · .json

Install

martinholovsky/claude-skills-generator/speech-to-text · repository language: Shell

generated, unverified - the skill's exact subdirectory could not be determined; check the repository on GitHub

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

What does Speech To Text do?

Speech To Text is a skill that converts spoken audio or voice recordings into written text using Faster Whisper technology. It handles real-time streaming, supports multiple languages, and processes audio locally with immediate deletion of voice data for privacy and security.

How do I convert audio to text with this skill?

Speech To Text accepts audio input and transcribes it into written text. The skill processes your audio locally on your device, ensuring immediate deletion of voice data after transcription. It supports various audio formats and can handle both pre-recorded files and real-time streaming.

Can Speech To Text transcribe speech for accessibility?

Yes, Speech To Text transcribes speech for accessibility and documentation purposes. By converting spoken words to text, it enables voice-based input workflows and creates written records of audio content, making information accessible in text form for users who need it.

What audio transcription tool features does Speech To Text offer?

Speech To Text provides hardware-optimized model selection, real-time streaming capabilities, and multi-language support. The skill prioritizes secure handling by processing audio locally and deleting voice information immediately after transcription, ensuring your data remains private.

How does Speech To Text handle my voice data?

Speech To Text processes all audio locally on your device rather than sending it to external servers. Voice data is deleted immediately after transcription completes, ensuring maximum privacy and security. The skill never stores or retains your audio recordings.

What license does Speech To Text use?

Speech To Text is released under the Unlicense, which places it in the public domain. This means you have complete freedom to use, modify, and distribute the skill without restrictions or attribution requirements.

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Convert spoken audio or voice recordings into written text”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

interview-transcription
by jamditis · jamditis/claude-skills-journalism

Convert interview recordings into searchable transcripts with word-level timestamps and speaker identification. Extract and verify quotes for publication, track sources, and organize multi-interview projects with built-in templates for manual transcription and quote management.

MITupdated Jul 2026
★ 342repo stars
Text To Speech
by martinholovsky · martinholovsky/claude-skills-generator

This skill provides expert-level text-to-speech implementation using Kokoro TTS, enabling real-time voice synthesis with customizable voices and prosody control. It emphasizes secure content handling, performance optimization through streaming and caching, and resource-efficient audio generation suitable for voice assistant applications.

Unlicenseupdated Dec 2025
★ 45repo stars
Mcp
by martinholovsky · martinholovsky/claude-skills-generator

Master MCP by building secure servers and clients that connect AI assistants to external tools and resources. This skill covers tool registration, transport configuration, authorization patterns, and performance optimization through test-driven development and security-first principles.

Unlicenseupdated Dec 2025
★ 45repo stars
Alibabacloud Bailian Voice Creator
by aliyun · aliyun/alibabacloud-aiops-skills

Convert text to natural-sounding speech or transcribe audio files using Alibaba Cloud's DashScope API. This skill handles both speech synthesis with customizable voice styles and speech recognition for audio up to 12 hours long, supporting 30+ languages and multiple audio formats.

no license declared → metadata onlyupdated Jul 2026
★ 198repo stars
Whisper
by graniet · graniet/kheish

Whisper is OpenAI's multilingual automatic speech recognition model, trained on 680,000 hours of audio data. It handles transcription, translation to English, and language identification across 99 languages with six configurable model sizes ranging from 39M to 1550M parameters. Use it for podcast transcription, meeting notes, noisy audio processing, or any speech-to-text task requiring robust multilingual support.

Apache-2.0updated Jul 2026
★ 264repo stars
Appsec Expert
by martinholovsky · martinholovsky/claude-skills-generator

Appsec Expert provides specialized guidance for securing applications throughout the development lifecycle. It covers threat modeling with STRIDE, vulnerability testing via SAST/DAST/SCA tools, secure coding patterns, and DevSecOps automation. Use it to identify vulnerabilities, design defense-in-depth controls, and remediate security issues with verified, production-ready approaches.

Unlicenseupdated Dec 2025
★ 45repo stars

More skills faster-whisper (MIT) · whisper-transcription (MIT) · Glsl (Unlicense)

Tags
audio-processingvoice-recognitiontranscription-engineaccessibility-toolspeech-conversionreal-time-transcriptionvoice-inputaudio-analysis