{"enrichment":{"faq":[{"a":"GPU Server Management covers driver installation across distributions. On Ubuntu, download drivers from NVIDIA's website or use package managers (apt), then disable Nouveau, install build tools, and run the installer. GPU Server Management also guides RHEL/CentOS setups. After installation, verify with `nvidia-smi` and enable persistence mode for production workloads.","q":"How do I install NVIDIA GPU drivers on Ubuntu?"},{"a":"GPU Server Management provides end-to-end LLM server provisioning. Start by installing NVIDIA drivers and CUDA toolkit, then configure Docker with the NVIDIA Container Toolkit for containerized inference. Set up multi-GPU topology with NVLink if available, enable MIG partitioning for multi-tenant isolation, and deploy monitoring with DCGM exporter and Prometheus for real-time metrics tracking.","q":"What's the best way to set up a GPU server for LLM?"},{"a":"GPU Server Management explains MIG setup for workload isolation. Enable MIG mode via nvidia-smi, partition GPUs into instances (1g, 2g, 3g profiles), and assign compute instances to applications. This allows multiple workloads to run safely in parallel. GPU Server Management covers instance configuration, memory allocation, and verification steps for production deployments.","q":"How do I configure MIG partitioning on A100/H100 GPUs?"},{"a":"GPU Server Management recommends using nvidia-smi for quick checks and DCGM exporter with Prometheus for production monitoring. nvidia-smi displays temperature, power draw, and utilization instantly. For persistent metrics, GPU Server Management guides deploying DCGM exporter to expose GPU health data to Prometheus, enabling dashboards and alerting on thermal throttling or performance degradation.","q":"How can I monitor GPU temperature and utilization in real-time?"},{"a":"GPU Server Management addresses common GPU failures and performance issues. XID errors indicate hardware or driver problems\u2014check logs, update drivers, and verify cooling. For thermal throttling, GPU Server Management recommends adjusting power limits with nvidia-smi, improving airflow, and monitoring DCGM metrics. Persistent issues may require hardware replacement or workload redistribution across servers.","q":"What should I do about GPU XID errors and thermal throttling?"},{"a":"GPU Server Management covers memory optimization techniques for inference. Disable ECC if not required to free VRAM, use mixed precision (FP16/INT8), enable persistence mode to reduce allocation overhead, and configure memory pooling in inference frameworks. GPU Server Management also guides multi-GPU load balancing and MIG partitioning to maximize throughput while minimizing latency on production inference clusters.","q":"How do I optimize GPU memory for inference servers?"}],"shadow_tags":["hardware-provisioning","inference-optimization","distributed-training","container-orchestration","performance-monitoring","resource-isolation","thermal-management","driver-compatibility"],"summary_rewrite":"GPU Server Management guides you through provisioning and configuring NVIDIA GPU infrastructure for LLM inference and model training. It covers driver and CUDA toolkit installation, Docker GPU integration, multi-GPU topology setup, and production monitoring with DCGM and Prometheus metrics."},"files":[{"bytes":7522,"path":"infrastructure/servers/gpu-server-management/SKILL.md","sha256":"760971b3b4e34f5fc05452ba1bfa6df83e0871d51606d5b96c7452d142eb3348","url":"https://skillfed.io/files/BagelHole/DevOps-Security-Agent-Skills/gpu-server-management/819fd25e/SKILL.md"}],"id":"BagelHole/DevOps-Security-Agent-Skills/gpu-server-management","links":{"html":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/gpu-server-management","md":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/gpu-server-management.md","repo":"https://github.com/BagelHole/DevOps-Security-Agent-Skills"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":4,"language":"Shell","last_updated":"2026-05-22","license":"MIT","name":"gpu-server-management","publisher":"BagelHole","stars":44},"relations":{"similar":[{"id":"BagelHole/DevOps-Security-Agent-Skills/gpu-kubernetes-operations"},{"id":"JosiahSiegel/claude-plugin-marketplace/ffmpeg-docker-containers"},{"id":"Mathews-Tom/armory/gpu-optimizer"},{"id":"aliyun/alibabacloud-aiops-skills/alibabacloud-ecs-gpu-diagnosis"},{"id":"agentsope/SkillAlchemy/agentsop-llm-engine-selection"},{"id":"BagelHole/DevOps-Security-Agent-Skills/llm-inference-scaling"},{"id":"agentsope/SkillAlchemy/agentsop-vllm"},{"id":"arpitg1304/robotics-agent-skills/docker-ros2-development"},{"id":"wanshuiyin/Auto-claude-code-research-in-sleep/system-profile"},{"id":"OpenLAIR/dr-claw/aris-system-profile"}]},"slug":{"owner":"BagelHole","repo":"DevOps-Security-Agent-Skills","skill":"gpu-server-management"},"version":"819fd25e"}
