{"categories":[{"label":"Monitoring","url":"https://skillfed.io/packages/category/system-monitoring/4"}],"enrichment":{"capability":"Monitors, debugs, and profiles workloads running on cloud accelerators (TPUs and GPUs), with built-in support for uploading diagnostic data to Vertex AI Tensorboard.","skillfed_tags":["accelerator-profiling","vertex-ai","google-cloud"],"use_cases":["Profile workloads running on TPUs and automatically upload traces to Vertex AI Tensorboard for analysis.","Monitor GPU training jobs and stream diagnostic logs to a centralized Tensorboard instance for debugging.","Set up a Vertex AI Experiment with automatic Tensorboard attachment to compare multiple accelerator runs.","Capture and persist accelerator performance metrics without writing custom Google Cloud API boilerplate.","Debug training failures by reviewing uploaded logs in Vertex AI Tensorboard after a job completes."],"what_it_does":"Cloud Accelerator Diagnostics is a library for monitoring and profiling machine learning workloads on cloud TPUs and GPUs. It wraps Vertex AI Tensorboard integration, allowing you to automatically capture and upload diagnostic logs from accelerator runs without manual instrumentation of your training code.\n\nThe package provides three main entry points: creating Vertex AI Tensorboard instances, creating Experiments within those instances, and starting a background thread that continuously monitors a log directory and uploads new data to Tensorboard. It is designed to work alongside profiling frameworks, and handles the Google Cloud authentication and API calls on your behalf. The main runtime dependency is google-cloud-aiplatform.","worth_installing":"Yes, if you are running workloads on Google Cloud TPUs or GPUs and want streamlined Tensorboard integration. The package is actively maintained, has low install friction, and eliminates boilerplate for Vertex AI setup. However, verify the license treatment before use in commercial contexts, and confirm that google-cloud-aiplatform's dependencies fit your environment. Not relevant for non-Google-Cloud accelerator setups."},"id":"cloud-accelerator-diagnostics","links":{"html":"https://skillfed.io/packages/cloud-accelerator-diagnostics","md":"https://skillfed.io/packages/cloud-accelerator-diagnostics.md","pypi":"https://pypi.org/project/cloud-accelerator-diagnostics/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2024-10-15","license_spdx":null,"license_treatment":"unclear","name":"cloud-accelerator-diagnostics","python_support":"supports_current","summary":"Monitor, debug and profile the jobs running on Cloud accelerators like TPUs and GPUs."},"popularity":{"monthly_downloads":274717,"position":8186,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.1.1"}
