$npx skillfedfor your agent

kugelaudio

Official Python SDK for KugelAudio TTS API

With conditionsPyPI SpeechReleased Aug 202678.6K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — kugelaudio-1.9.0-py3-none-any.whl
v1.9.0 · released 2026-08-07 · Python >=3.9 · 2 runtime deps: httpx, websockets

Yes, if you need a KugelAudio TTS client or local turn detection for conversational AI. The core SDK is lightweight with minimal dependencies and active maintenance. The optional turn-detection extra adds real value for voice applications but requires Python 3.11+, Hugging Face access, and roughly 1.55 GiB of model runtime. No known security vulnerabilities. MIT license imposes no restrictions.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a valid KugelAudio API key.
  • Turn detection extra requires Python 3.11 or newer and Hugging Face authentication for the private model repository.
  • Low friction: pure Python wheel with only httpx and websockets as runtime dependencies.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it straightforward to integrate into proprietary applications.

last release 2026-08-07 (7 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 78,555 downloads/mo, #14,431 on PyPI

Verify before relying

pip install kugelaudio

from kugelaudio import KugelAudio

client = KugelAudio(api_key="your_api_key")
audio = client.tts.generate(
    text="Hello, world!",
    model_id="kugel-1-turbo",
)
audio.save("output.wav")
  • Whether the turn-detection model bundle size and download time are acceptable for typical deployment workflows.
  • Performance characteristics and latency of turn detection on various hardware configurations.
  • Availability and stability of the Hugging Face model repository for production use.
Same gist for agents: .md · .json

What it is and what it does

KugelAudio is the official Python client for the KugelAudio text-to-speech API, handling speech synthesis via WebSocket streaming. The core SDK requires only httpx and websockets, making it lightweight to install. It provides a simple interface to generate audio from text using specified model IDs and save the output to files.

An optional turn-detection extra adds local, CPU-based conversation endpoint detection using ONNX Runtime and PyTorch. This feature downloads a version-pinned model bundle from Hugging Face, verifies SHA-256 checksums, and runs inference locally without network calls after the initial model cache. It's designed for latency-sensitive voice applications—particularly LiveKit Agents integrations—where you need to detect when a user has finished speaking. The turn detector supports multiple languages and provides detailed decision reasoning, including incomplete-turn timeouts and barge-in handling for overlapping speech.

Use it for

  • Generate speech from text in a web or mobile backend, streaming audio over WebSocket to reduce latency.
  • Build a conversational voice agent that detects when users finish speaking using local turn detection without API round-trips.
  • Integrate KugelAudio TTS into a LiveKit Agents application with manual turn handling and configurable silence thresholds.
  • Implement a voice interface where you need to balance synthesis quality with low-latency streaming and local endpoint detection.
  • Run turn detection offline after caching the model, suitable for edge deployments or privacy-sensitive applications.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need a KugelAudio TTS client or local turn detection for conversational AI.

The core SDK is lightweight with minimal dependencies and active maintenance. The optional turn-detection extra adds real value for voice applications but requires Python 3.11+, Hugging Face access, and roughly 1.55 GiB of model runtime. No known security vulnerabilities. MIT license imposes no restrictions.

Install

kugelaudio on PyPI

Before you install

Low friction: pure Python wheel with only httpx and websockets as runtime dependencies. Actively maintained with a release within the past week. Optional turn-detection extra requires Python 3.11+ and downloads a pinned ONNX model bundle from Hugging Face.

Requires a valid KugelAudio API key. Turn detection extra requires Python 3.11 or newer and Hugging Face authentication for the private model repository.

License in practice

MIT license permits commercial and private use with minimal restrictions, making it straightforward to integrate into proprietary applications.

Quickstart

pip install kugelaudio

from kugelaudio import KugelAudio

client = KugelAudio(api_key="your_api_key")
audio = client.tts.generate(
    text="Hello, world!",
    model_id="kugel-1-turbo",
)
audio.save("output.wav")

Verify before relying

  • Whether the turn-detection model bundle size and download time are acceptable for typical deployment workflows.
  • Performance characteristics and latency of turn detection on various hardware configurations.
  • Availability and stability of the Hugging Face model repository for production use.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
httpxwebsockets
MaintenanceActively maintained 7 days since the last release
First released
Downloads78,555 / month, #14,431 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Multimedia :: Sound/Audio :: Speech

Evidence: kugelaudio-1.9.0-py3-none-any.whl

Tags

Capabilities
text-to-speech api clienttts python sdkturn detection speechconversational ai endpointwebsocket audio streamingvoice synthesis apispeech endpoint detection
Topics
tts-apiturn-detectionlivekit-integration
PyPI keywords
audiostreamingtext-to-speechttswebsocket

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “turn detection speech”

Give your agent the search over MCP, or paste the wish link into any chat.

More Speech packages

SpeechRecognition Worth it
PyPI · Python Modules · released Jun 2026

Performs speech recognition and transcription using multiple online and offline engines, including Google, OpenAI Whisper, CMU Sphinx, and others.

The main gotcha is that most engines require optional dependencies or API credentials.

BSD-3-Clausepure Python · 3.9+
11.7Mdownloads / mo
gTTS With conditions
PyPI · Libraries · released Nov 2024

gTTS converts text to speech using Google Translate's API, writing MP3 audio to files, file-like objects, or stdout via Python library or command-line tool.

However, be aware that it depends on Google Translate's undocumented API—upstream changes can break it without notice, and it is not a substitute for official Google…

MITpure Python · 3.7+
5.3Mdownloads / mo
lhotse Worth it
PyPI · Python Modules · released Apr 2026

Lhotse prepares multimodal (speech, audio, video, image, text) data for machine learning model training with flexible pipelines, on-the-fly augmentation, and efficient data loading.

Install it if you are building speech, audio, or multimodal training pipelines; skip it if you only need simple audio I/O without data augmentation or complex dataset…

Apache-2.0pure Python · 3.8.0+
1.3Mdownloads / mo
piper-tts With conditions
PyPI · Speech · released Aug 2026

Piper TTS is a local neural text-to-speech engine that converts text to speech using embedded phonemization, with support for multiple languages and voices.

copyleftcompiled wheel · 3.9+
891.3Kdownloads / mo
funasr Worth it
PyPI · Python Modules · released Aug 2026

FunASR is a speech recognition toolkit that transcribes audio offline or via streaming, with integrated voice activity detection, speaker identification, punctuation restoration, and emotion/audio-event tagging across multiple languages and deployment targets.

Install it if you need speaker diarization, emotion detection, streaming support, or self-hosted deployment.

MITpure Python · 3.7.0+
497.9Kdownloads / mo
pocketsphinx With conditions
PyPI · Speech · released Jun 2026

PocketSphinx provides Python bindings for Carnegie Mellon University's open-source speech recognition engine, enabling continuous speech-to-text and keyword spotting from live microphone input or audio files.

BSD-3-Clausecompiled wheel
382.6Kdownloads / mo

See also livekit-plugins-turn-detector · cartesia · livekit-plugins-soniox · livekit-plugins-speechmatics · livekit-plugins-inworld · aic-sdk · soniox · speechmatics-voice · livekit-plugins-cartesia · rev-ai