$npx skillfedfor your agent

gpu-kubernetes-operations

Deploy and operate production-grade GPU clusters in Kubernetes with built-in support for NVIDIA device plugins, MIG partitioning, and time-slicing. Monitor GPU health via DCGM metrics and Prometheus, configure autoscaling policies, and troubleshoot scheduling and driver issues across your AI infrastructure.

GPU Kubernetes Operations helps you set up and manage GPU-backed Kubernetes clusters for AI inference and training workloads.

AI-generated summary based on this skill's SKILL.md

44 4 MITupdated by BagelHole

Decision gist · record as of 2026-05-22

GPU Kubernetes Operations helps you set up and manage GPU-backed Kubernetes clusters for AI inference and training workloads. Deploy and operate production-grade GPU clusters in Kubernetes with built-in support for NVIDIA device plugins, MIG partitioning, and time-slicing. Monitor GPU health via DCGM metrics and Prometheus, configure autoscaling policies, and troubleshoot scheduling and driver issues across your AI infrastructure.

manual: git clone https://github.com/BagelHole/DevOps-Security-Agent-Skills → cp -r DevOps-Security-Agent-Skills/infrastructure/local-ai/gpu-kubernetes-operations ~/.claude/skills/gpu-kubernetes-operations
infrastructure/local-ai/gpu-kubernetes-operations/SKILL.md · version b341c97d

Use it when

  • gpu-kubernetes-operations supports Multi-Instance GPU (MIG) partitioning for A100 and H100 GPUs.
  • gpu-kubernetes-operations integrates DCGM (Data Center GPU Manager) for comprehensive GPU health monitoring.

Verify before relying

Read SKILL.md below before installing (1 file). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

BagelHole/DevOps-Security-Agent-Skills/gpu-kubernetes-operations · repository language: Shell

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

How to set up GPU nodes in Kubernetes?

gpu-kubernetes-operations enables GPU node setup through NVIDIA device plugins and the GPU Operator. Install the NVIDIA GPU Operator on your cluster to automatically deploy drivers, container toolkit, and device plugins across GPU nodes. Configure node pools with GPU labels and taints, then schedule AI workloads using GPU resource requests (nvidia.com/gpu). The skill covers full lifecycle management from driver installation through pod scheduling.

What is MIG partitioning and how do I use it in Kubernetes?

gpu-kubernetes-operations supports Multi-Instance GPU (MIG) partitioning for A100 and H100 GPUs, allowing you to divide a single GPU into multiple isolated instances. Enable MIG mode on your GPU nodes, partition GPUs into compute instances, and configure the device plugin to expose partitions as separate resources. This maximizes utilization by running multiple smaller workloads simultaneously on one physical GPU.

How can I monitor GPU health and metrics with DCGM in Kubernetes?

gpu-kubernetes-operations integrates DCGM (Data Center GPU Manager) for comprehensive GPU health monitoring. Deploy DCGM exporter as a DaemonSet to collect metrics like temperature, power, memory usage, and error counts. Scrape metrics into Prometheus and set up alerts for GPU failures, thermal issues, and XID errors. This enables proactive detection of hardware problems before they impact your AI workloads.

How do I troubleshoot GPU pod pending issues in Kubernetes?

gpu-kubernetes-operations provides troubleshooting guidance for GPU scheduling failures. Check device plugin logs, verify GPU resource availability with `kubectl describe nodes`, and confirm pod requests match available GPU types. Common causes include driver mismatches, insufficient GPU memory, node taints without matching tolerations, or device plugin crashes. Use DCGM health checks to rule out hardware failures.

What GPU time-slicing setup does this skill support?

gpu-kubernetes-operations enables GPU time-slicing to share a single GPU across multiple pods sequentially. Configure the device plugin with sharing mode, set time-slice durations, and define fairness policies. This allows cost-effective multi-tenant AI inference clusters where pods take turns accessing GPU compute. Combine with MIG partitioning for fine-grained resource isolation and higher throughput.

How can I implement GPU autoscaling and cost optimization?

gpu-kubernetes-operations supports GPU autoscaling through Kubernetes HPA and cluster autoscaler integration. Define GPU resource requests in pod specs, set HPA scaling policies based on GPU utilization metrics, and configure node auto-provisioning to add GPU nodes on demand. Combine with time-slicing and MIG to maximize utilization, reducing idle GPU costs while maintaining performance for your AI inference and training workloads.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

GPU Kubernetes Operations

Run resilient and cost-efficient GPU clusters for production AI workloads.

When to Use This Skill

  • Setting up GPU node pools in Kubernetes for AI inference or training
  • Configuring NVIDIA device plugin and GPU operator
  • Implementing MIG partitioning to share GPUs across workloads
  • Building GPU-aware autoscaling policies
  • Monitoring GPU health with DCGM and Prometheus
  • Troubleshooting GPU scheduling, driver, or OOM issues

Prerequisites

  • Kubernetes 1.28+ cluster with GPU-capable nodes
  • NVIDIA GPUs (A10, L4, A100, H100, or similar)
  • NVIDIA drivers installed on nodes (535+ recommended)
  • Helm 3 for operator and plugin installation
  • Prometheus stack for metrics collection

NVIDIA GPU Operator Installation

The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.

```bash

Add NVIDIA Helm repo

helm repo add nvidia

(truncated - see the full file via the links below)

File tree — 1 file
infrastructure/local-ai/gpu-kubernetes-operations/SKILL.md

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Set up and manage GPU-backed Kubernetes clusters for AI workloads”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

gpu-server-management
by BagelHole · BagelHole/DevOps-Security-Agent-Skills

GPU Server Management guides you through provisioning and configuring NVIDIA GPU infrastructure for LLM inference and model training. It covers driver and CUDA toolkit installation, Docker GPU integration, multi-GPU topology setup, and production monitoring with DCGM and Prometheus metrics.

MITupdated May 2026
★ 44repo stars
llm-inference-scaling
by BagelHole · BagelHole/DevOps-Security-Agent-Skills

This skill enables dynamic scaling of LLM inference workloads across Kubernetes clusters using KEDA and Prometheus metrics tied to GPU utilization and request queues. It covers vLLM deployment, queue-based job scaling with Redis, spot instance strategies, and cluster autoscaler configuration to handle traffic spikes while optimizing costs.

MITupdated May 2026
★ 44repo stars
llmops-platform-engineering
by BagelHole · BagelHole/DevOps-Security-Agent-Skills

LLMOps Platform Engineering teaches you to architect internal LLM platforms that balance rapid experimentation with production safety. You'll implement model promotion pipelines with automated quality and safety gates, canary validation, and rollback capabilities, plus set up A/B testing infrastructure and observability across Kubernetes and cloud inference.

MITupdated May 2026
★ 44repo stars
model-serving-kubernetes
by BagelHole · BagelHole/DevOps-Security-Agent-Skills

Run production ML inference on Kubernetes using KServe or NVIDIA Triton, with built-in support for canary traffic splitting, request-based autoscaling, and GPU resource allocation. The skill covers model versioning, A/B testing patterns, and dynamic batching for throughput optimization.

MITupdated May 2026
★ 44repo stars
aks-automatic-2025
by JosiahSiegel · JosiahSiegel/claude-plugin-marketplace

AKS Automatic 2025 is a fully-managed Kubernetes offering that handles cluster operations, security patching, and node provisioning automatically. It includes Karpenter-based dynamic scaling, Microsoft Entra integration, Azure CNI Overlay networking with Cilium, and built-in monitoring through Azure Monitor. Use this skill to deploy production clusters, configure autoscaling with HPA/VPA/KEDA, set up workload identity, and understand the new billing model.

MITupdated Jun 2026
★ 49repo stars
cis-benchmarks
by BagelHole · BagelHole/DevOps-Security-Agent-Skills

This skill automates CIS benchmark auditing across Linux and Kubernetes environments using industry-standard tools. Run security assessments, identify compliance gaps, and track remediation through a structured workflow that includes scanning, analysis, fixes, and validation.

MITupdated May 2026
★ 44repo stars

More skills azure-aks (MIT)

Tags
gpu-resource-managementcontainer-orchestrationml-infrastructurecluster-scalinghardware-monitoringworkload-schedulingcost-efficiencyfault-toleranceperformance-tuning