cloud-tpu-diagnostics
Monitor, debug and profile the jobs running on Cloud TPU.
Decision gist · record as of 2026-08-14
Yes, if you run workloads on Cloud TPU and need built-in fault and hang diagnostics. The package is actively maintained, has no external dependencies, and integrates directly with Google Cloud Logging. However, verify the license terms before production use, and confirm it meets your TPU environment's requirements.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Designed for use on Cloud TPU VMs; stack trace collection and cloud upload features require a TPU environment and Google Cloud credentials.
- Low friction installation as a pure Python wheel with no runtime dependencies.
- Actively maintained as of 2026-04-08 with recent updates, though the project remains relatively young (first released 2023-06-07).
License · maintenance · safety
(unclear) — License status is unclear—the package metadata lists no SPDX identifier or raw license text. Before production use, verify the actual license terms in the repository.
last release 2023-12-08 (980 days) · last repo commit 2026-04-08 · 17 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 191,100 downloads/mo, #9,891 on PyPI
Alternatives
Verify before relying
from cloud_tpu_diagnostics import diagnostic
from cloud_tpu_diagnostics.configuration import stack_trace_configuration, diagnostic_configuration, debug_configuration
stack_trace_config = stack_trace_configuration.StackTraceConfig(collect_stack_trace=True, stack_trace_to_cloud=True)
debug_config = debug_configuration.DebugConfig(stack_trace_config=stack_trace_config)
diagnostic_config = diagnostic_configuration.DiagnosticConfig(debug_config=debug_config)
with diagnostic.diagnose(diagnostic_config):
run_job()- Exact scope of what 'monitor' and 'profile' capabilities include beyond stack trace collection.
- Whether the package works on non-TPU systems or is strictly TPU-only.
- Performance overhead of periodic stack trace collection at different intervals.
What it is and what it does
Cloud TPU Diagnostics is a debugging and monitoring library for workloads running on Google Cloud TPU. It captures Python stack traces when faults occur (segmentation faults, floating-point exceptions, illegal operations) and can periodically collect snapshots to identify where a job is hung. Stack traces can be displayed on the console or uploaded to Google Cloud Logging for centralized troubleshooting.
The package is configured through a composition of configuration objects: you define stack trace behavior (whether to collect, where to send them, and how often), wrap that in a debug configuration, then wrap that in a diagnostic configuration, and finally use a context manager around the code you want to monitor. It has no external runtime dependencies and installs as a pure Python wheel, making it straightforward to add to a TPU VM.
Use it for
- Capture stack traces when a TPU job crashes with a segmentation fault or other signal to diagnose the root cause.
- Periodically snapshot the call stack of a hung TPU workload to identify where it is stuck.
- Upload diagnostic traces to Google Cloud Logging for centralized analysis and troubleshooting across multiple TPU jobs.
- Debug training or inference jobs on Cloud TPU by collecting stack traces at configurable intervals (default 10 minutes).
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you run workloads on Cloud TPU and need built-in fault and hang diagnostics.
The package is actively maintained, has no external dependencies, and integrates directly with Google Cloud Logging. However, verify the license terms before production use, and confirm it meets your TPU environment's requirements.
Install
cloud-tpu-diagnostics on PyPI
Before you install
Low friction installation as a pure Python wheel with no runtime dependencies. Actively maintained as of 2026-04-08 with recent updates, though the project remains relatively young (first released 2023-06-07).
Designed for use on Cloud TPU VMs; stack trace collection and cloud upload features require a TPU environment and Google Cloud credentials.
License in practice
License status is unclear—the package metadata lists no SPDX identifier or raw license text. Before production use, verify the actual license terms in the repository.
Quickstart
from cloud_tpu_diagnostics import diagnostic
from cloud_tpu_diagnostics.configuration import stack_trace_configuration, diagnostic_configuration, debug_configuration
stack_trace_config = stack_trace_configuration.StackTraceConfig(collect_stack_trace=True, stack_trace_to_cloud=True)
debug_config = debug_configuration.DebugConfig(stack_trace_config=stack_trace_config)
diagnostic_config = diagnostic_configuration.DiagnosticConfig(debug_config=debug_config)
with diagnostic.diagnose(diagnostic_config):
run_job()
Verify before relying
- Exact scope of what 'monitor' and 'profile' capabilities include beyond stack trace collection.
- Whether the package works on non-TPU systems or is strictly TPU-only.
- Performance overhead of periodic stack trace collection at different intervals.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 980 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 191,100 / month, #9,891 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: cloud_tpu_diagnostics-0.1.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “cloud tpu debugging”
- cloud-tpu-diagnosticsCollects stack traces and diagnostic data from jobs running on Cloud…
- tpu-infoCLI tool that detects Cloud TPU devices and reads runtime metrics…
- cloud-accelerator-diagnosticsMonitors, debugs, and profiles workloads running on cloud…
Give your agent the search over MCP, or paste the wish link into any chat.
More Monitoring packages
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
Provides generated Python code for OpenTelemetry semantic conventions, enabling standardized attribute naming and constant definitions for instrumentation and telemetry collection.
Install it if you are using OpenTelemetry and want to follow semantic conventions correctly.
Provides the reference implementation of the OpenTelemetry API for collecting and exporting traces, metrics, and logs from Python applications.
Provides the abstract API and interfaces for OpenTelemetry instrumentation in Python, defining how to emit traces, metrics, and logs without tying code to a specific SDK implementation.
Exports OpenTelemetry observability data to an OpenTelemetry Collector using Protobuf-encoded messages over HTTP.
Install it if you are using OpenTelemetry in Python and need to send data to a Collector over HTTP.
Provides automatic instrumentation commands and programmatic APIs to inject distributed tracing into Python applications without code changes, detecting and instrumenting packages used by your program.
Install it if you need distributed tracing without code changes and have compatible instrumented packages in your environment.
See also google-cloud-tpu · google-cloud-profiler · cloud-accelerator-diagnostics · tpu-info · slurm-usage · opentelemetry-exporter-gcp-trace · google-cloud-mldiagnostics · osprofiler · libtpu · oslo.reports