{"enrichment":{"faq":[{"a":"llm-inference-scaling provides KEDA-based autoscaling that monitors GPU utilization and request queue depth to dynamically scale LLM inference pods. It integrates Prometheus metrics collection with vLLM deployments, allowing the system to add or remove pod replicas based on real-time GPU load and incoming API traffic patterns, ensuring efficient resource use across your Kubernetes cluster.","q":"How does llm-inference-scaling enable auto-scale LLM pods on Kubernetes?"},{"a":"llm-inference-scaling uses KEDA to connect custom Prometheus metrics\u2014including GPU utilization and Redis queue length\u2014to Kubernetes' horizontal pod autoscaler. KEDA acts as the bridge between your monitoring stack and scaling decisions, enabling GPU-aware triggers that traditional HPA cannot achieve, so LLM workloads scale intelligently based on actual accelerator demand.","q":"What role does KEDA play in GPU-aware scaling with this skill?"},{"a":"Yes. llm-inference-scaling includes strategies for deploying LLM inference on spot instances and preemptible GPUs within Kubernetes, combined with cluster autoscaler configuration. By pairing queue-based scaling with cost-aware node provisioning, it minimizes expenses during traffic lulls while maintaining performance during spikes, reducing overall LLM serving infrastructure costs.","q":"Can llm-inference-scaling reduce LLM serving costs using spot instances?"},{"a":"llm-inference-scaling implements queue-based autoscaling using Redis or similar systems to buffer incoming requests. KEDA monitors queue depth via Prometheus and triggers pod scaling before requests timeout. This decouples request arrival from processing capacity, allowing the system to absorb traffic spikes gracefully and scale compute resources to match queue buildup.","q":"How does llm-inference-scaling handle unpredictable LLM API traffic?"},{"a":"llm-inference-scaling relies on Prometheus metrics including GPU utilization, GPU memory usage, request queue length (from Redis), and vLLM-specific counters like pending requests. These metrics feed KEDA scalers, enabling multi-dimensional scaling logic that reacts to both resource saturation and workload backlog, not just CPU or memory.","q":"What metrics does llm-inference-scaling use for scaling decisions?"},{"a":"llm-inference-scaling can orchestrate multi-model LLM serving by scaling separate deployment groups per model while sharing GPU node pools. Each model's queue and GPU metrics are tracked independently in Prometheus, allowing KEDA to balance scaling decisions across models and maximize shared GPU resource utilization on your Kubernetes cluster.","q":"Does llm-inference-scaling support managing multiple LLM models?"}],"shadow_tags":["gpu-orchestration","event-driven-scaling","cost-optimization","queue-management","inference-serving","spot-instances","kubernetes-autoscaling","llm-deployment","resource-efficiency","multi-model-serving"],"summary_rewrite":"This skill enables dynamic scaling of LLM inference workloads across Kubernetes clusters using KEDA and Prometheus metrics tied to GPU utilization and request queues. It covers vLLM deployment, queue-based job scaling with Redis, spot instance strategies, and cluster autoscaler configuration to handle traffic spikes while optimizing costs."},"files":[{"bytes":7912,"path":"infrastructure/local-ai/llm-inference-scaling/SKILL.md","sha256":"d195f1a7410bcddbc48edf0ce170f94a69796de194b2a07f29a08a02d481510e","url":"https://skillfed.io/files/BagelHole/DevOps-Security-Agent-Skills/llm-inference-scaling/720f5837/SKILL.md"}],"id":"BagelHole/DevOps-Security-Agent-Skills/llm-inference-scaling","links":{"html":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/llm-inference-scaling","md":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/llm-inference-scaling.md","repo":"https://github.com/BagelHole/DevOps-Security-Agent-Skills"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":4,"language":"Shell","last_updated":"2026-05-22","license":"MIT","name":"llm-inference-scaling","publisher":"BagelHole","stars":44},"relations":{"similar":[{"id":"BagelHole/DevOps-Security-Agent-Skills/model-serving-kubernetes"},{"id":"BagelHole/DevOps-Security-Agent-Skills/gpu-kubernetes-operations"},{"id":"JosiahSiegel/claude-plugin-marketplace/aks-automatic-2025"},{"id":"ancoleman/ai-design-components/model-serving"},{"id":"rohitg00/kubectl-mcp-server/k8s-autoscaling"},{"id":"ancoleman/ai-design-components/operating-kubernetes"},{"id":"BagelHole/DevOps-Security-Agent-Skills/multi-tenant-llm-hosting"},{"id":"google/skills/gke-basics"},{"id":"Tomlord1122/tomtom-skill/cloud-architect"},{"id":"BagelHole/DevOps-Security-Agent-Skills/llmops-platform-engineering"}]},"slug":{"owner":"BagelHole","repo":"DevOps-Security-Agent-Skills","skill":"llm-inference-scaling"},"version":"720f5837"}
