google-cloud-speech
Google Cloud Speech API client library
Decision gist · record as of 2026-08-14
Yes, if you need speech-to-text in a Python application and have access to Google Cloud infrastructure. The library is production-stable, actively maintained, permissively licensed, and has low install friction. Requires Google Cloud project setup and billing; not suitable for offline-only or non-Google-Cloud deployments.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Google Cloud project setup, billing enabled, Speech-to-Text API enabled, and valid authentication credentials (typically via GOOGLE_APPLICATION_CREDENTIALS environment variable).
- Low friction installation with a pure-Python wheel.
- Actively maintained with a recent release (72 days old) and strong repository health (5371 stars).
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing use in commercial and private projects with minimal restrictions beyond attribution and liability disclaimers.
last release 2026-06-03 (72 days) · last repo commit 2026-08-14 · 5,371 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 122,581,480 downloads/mo, #300 on PyPI
Alternatives
Verify before relying
pip install google-cloud-speech
from google.cloud import speech
client = speech.SpeechClient()
config = speech.RecognitionConfig(language_code="en-US")
audio = speech.RecognitionAudio(uri="gs://bucket/audio.wav")
response = client.recognize(config=config, audio=audio)- Whether the package handles streaming audio transcription or only batch/file-based recognition.
- Supported audio formats, sample rates, and language coverage beyond the en-US example.
- Pricing model and cost implications for typical usage patterns.
- Whether real-time transcription latency meets specific application requirements.
What it is and what it does
This is the official Python client library for Google Cloud Speech-to-Text, a managed service that transcribes audio to text using Google's neural network models. It wraps the Google Cloud Speech API, allowing you to send audio (from files, streams, or Cloud Storage) and receive transcribed text with optional metadata like confidence scores and alternative interpretations.
The library handles authentication, request formatting, and response parsing through standard Google Cloud dependencies (google-api-core, google-auth, grpcio, proto-plus, protobuf). It's designed for developers building applications that need speech recognition—transcription services, voice commands, accessibility features, or audio analysis. Setup requires a Google Cloud project with billing enabled and the Speech-to-Text API activated; authentication is typically managed via environment variables or application default credentials.
Use it for
- Transcribe audio files stored in Google Cloud Storage or local filesystem for document creation or archival.
- Build voice-command interfaces that convert user speech input to actionable text in real-time applications.
- Add accessibility features to applications by automatically generating captions or transcripts from audio content.
- Analyze customer support calls or meeting recordings by extracting and processing spoken content as text.
- Implement multilingual transcription pipelines for international applications or content processing workflows.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need speech-to-text in a Python application and have access to Google Cloud infrastructure.
The library is production-stable, actively maintained, permissively licensed, and has low install friction. Requires Google Cloud project setup and billing; not suitable for offline-only or non-Google-Cloud deployments.
Install
google-cloud-speech on PyPI
Before you install
Low friction installation with a pure-Python wheel. Actively maintained with a recent release (72 days old) and strong repository health (5371 stars). Supports modern Python versions (3.10+) and depends on standard Google Cloud libraries.
Requires Google Cloud project setup, billing enabled, Speech-to-Text API enabled, and valid authentication credentials (typically via GOOGLE_APPLICATION_CREDENTIALS environment variable).
License in practice
Licensed under Apache-2.0 (permissive), allowing use in commercial and private projects with minimal restrictions beyond attribution and liability disclaimers.
Quickstart
pip install google-cloud-speech
from google.cloud import speech
client = speech.SpeechClient()
config = speech.RecognitionConfig(language_code="en-US")
audio = speech.RecognitionAudio(uri="gs://bucket/audio.wav")
response = client.recognize(config=config, audio=audio)
Verify before relying
- Whether the package handles streaming audio transcription or only batch/file-based recognition.
- Supported audio formats, sample rates, and language coverage beyond the en-US example.
- Pricing model and cost implications for typical usage patterns.
- Whether real-time transcription latency meets specific application requirements.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesgoogle-api-coregoogle-authgrpcioproto-plusprotobuf |
| Maintenance | Actively maintained 72 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 122,581,480 / month, #300 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet |
Evidence: google_cloud_speech-2.40.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “audio transcription google cloud”
- google-cloud-speechProvides a Python client for Google Cloud Speech-to-Text, converting…
- livekit-plugins-googleIntegrates Google Cloud AI services (Gemini, Speech-to-Text,…
- SpeechRecognitionPerforms speech recognition and transcription using multiple online…
Give your agent the search over MCP, or paste the wish link into any chat.
More Internet packages
Botocore provides low-level, data-driven access to Amazon Web Services APIs, serving as the foundation for the AWS CLI and boto3 libraries.
Install it if you need programmatic access to AWS services.
Provides an async client for AWS services using botocore and aiohttp, allowing you to call AWS APIs asynchronously within asyncio-based applications.
Install it if you need to call AWS services from async Python code; it is the standard way to do so.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides a platform-independent file locking mechanism to coordinate access to files across processes and threads.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides common Protocol Buffer message definitions used across Google Cloud APIs, enabling Python clients to interact with Google services.
See also google-cloud-texttospeech · azure-cognitiveservices-speech · google-cloud-translate · google-cloud-videointelligence · google-cloud-language · google-cloud-video-transcoder · google-cloud-automl · google-cloud-build · gTTS · google-cloud-appengine-logging