tpu-info
CLI tool to view TPU metrics
Decision gist · record as of 2026-08-14
Yes, if you work with Cloud TPU and need visibility into runtime metrics and device diagnostics. The tool is actively maintained, has no known vulnerabilities, and low install friction. It is purpose-built for TPU environments and requires an active workload to be useful; install it only if you have TPU hardware and supported ML frameworks available.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires an active TPU workload running a supported ML framework (JAX or PyTorch/XLA) on the same machine to access libtpu utilization metrics; without a workload, device detection may still work but runtime metrics will not be available.
- Low install friction with a pure-Python wheel distribution.
- Active maintenance with a recent release; last commit was 39 days ago.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions; you must include a copy of the license and state any material changes.
last release 2026-07-06 (39 days) · last repo commit 2026-08-14 · 32 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 270,511 downloads/mo, #8,234 on PyPI
Alternatives
Verify before relying
pip install tpu-info
# Then run the CLI command in a terminal where a TPU workload is active:
$ tpu-info
# or for streaming mode:
$ tpu-info --streaming --rate 2- Whether the package works on all Python versions from 3.8 onwards or if there are known incompatibilities (e.g., Python 3.12+ environment note in docs suggests possible limitations).
- Whether Pygrain and Orbax metrics support (added in 0.14.2) requires those frameworks to be installed or if they are optional integrations.
What it is and what it does
tpu-info is a command-line diagnostic tool for Google Cloud TPU environments that queries the libtpu runtime to expose hardware and performance metrics. It runs on machines with attached TPU devices and requires an active ML workload (JAX or PyTorch/XLA) to access full utilization data; without a workload, it can still detect TPU devices but will not report runtime metrics.
The tool offers two modes: a static snapshot mode that prints current metrics once, and a streaming mode that refreshes and displays metrics continuously at a configurable interval. Metrics include HBM memory usage, duty cycle, TensorCore utilization, buffer transfer latencies, host compute latency, and gRPC TCP performance. Recent versions added support for Pygrain and Orbax performance metrics from the local Prometheus server.
Use it for
- Monitor TPU memory and compute utilization in real time while running ML training or inference workloads to identify bottlenecks.
- Capture a one-time snapshot of TPU metrics for debugging or performance analysis of a specific computation.
- Diagnose TPU device availability and libtpu version compatibility before launching a workload.
- Track buffer transfer and gRPC latencies to optimize data pipeline performance on TPU clusters.
- Stream continuous metrics during development to watch TensorCore utilization and duty cycle as you iterate on model code.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with Cloud TPU and need visibility into runtime metrics and device diagnostics.
The tool is actively maintained, has no known vulnerabilities, and low install friction. It is purpose-built for TPU environments and requires an active workload to be useful; install it only if you have TPU hardware and supported ML frameworks available.
Install
tpu-info on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release; last commit was 39 days ago. Six runtime dependencies are all stable, widely-used libraries.
Requires an active TPU workload running a supported ML framework (JAX or PyTorch/XLA) on the same machine to access libtpu utilization metrics; without a workload, device detection may still work but runtime metrics will not be available.
License in practice
Apache-2.0 permissive license allows use in commercial and proprietary projects with minimal restrictions; you must include a copy of the license and state any material changes.
Quickstart
pip install tpu-info
# Then run the CLI command in a terminal where a TPU workload is active:
$ tpu-info
# or for streaming mode:
$ tpu-info --streaming --rate 2
Verify before relying
- Whether the package works on all Python versions from 3.8 onwards or if there are known incompatibilities (e.g., Python 3.12+ environment note in docs suggests possible limitations).
- Whether Pygrain and Orbax metrics support (added in 0.14.2) requires those frameworks to be installed or if they are optional integrations.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesgrpcioimmutabledictpackagingprometheus-clientprotobufrich |
| Maintenance | Actively maintained 39 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 270,511 / month, #8,234 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: tpu_info-0.14.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “TPU metrics monitoring”
- tpu-infoCLI tool that detects Cloud TPU devices and reads runtime metrics…
- google-cloud-tpuProvides a Python client library to interact with Google Cloud TPU…
- google-cloud-mldiagnosticsCollects metrics, configs, and performance profiles from ML workloads…
Give your agent the search over MCP, or paste the wish link into any chat.
More Monitoring packages
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
Provides generated Python code for OpenTelemetry semantic conventions, enabling standardized attribute naming and constant definitions for instrumentation and telemetry collection.
Install it if you are using OpenTelemetry and want to follow semantic conventions correctly.
Provides the reference implementation of the OpenTelemetry API for collecting and exporting traces, metrics, and logs from Python applications.
Provides the abstract API and interfaces for OpenTelemetry instrumentation in Python, defining how to emit traces, metrics, and logs without tying code to a specific SDK implementation.
Exports OpenTelemetry observability data to an OpenTelemetry Collector using Protobuf-encoded messages over HTTP.
Install it if you are using OpenTelemetry in Python and need to send data to a Collector over HTTP.
Provides automatic instrumentation commands and programmatic APIs to inject distributed tracing into Python applications without code changes, detecting and instrumenting packages used by your program.
Install it if you need distributed tracing without code changes and have compatible instrumented packages in your environment.
See also libtpu · cloud-tpu-diagnostics · tpu-inference · slurm-usage · google-cloud-mldiagnostics · google-cloud-tpu · pathwaysutils · cloud-accelerator-diagnostics · snowflake-cli-labs · tokamax