{"enrichment":{"faq":[{"a":"vllm-server helps you deploy vLLM\u2014a high-performance open-source LLM serving engine\u2014for production workloads. Start by installing vLLM, then configure your model, GPU allocation, and batch settings. Use Docker for containerized deployment or Kubernetes for orchestrated multi-node setups. vllm-server guides you through setting up continuous batching and paged attention to maximize throughput, configuring tensor parallelism across multiple GPUs, and exposing OpenAI-compatible API endpoints for seamless integration with existing applications.","q":"How do I deploy vllm server for production LLM inference?"},{"a":"vllm-server covers tensor parallelism configuration to distribute model inference across multiple GPUs. Enable tensor parallelism by specifying the number of GPUs and partition strategy in your vLLM launch command. vllm-server provides step-by-step guidance on configuring GPU allocation, setting tensor-parallel size, and validating that your model shards correctly across devices. This approach significantly reduces per-GPU memory requirements and enables serving larger models while maintaining low latency.","q":"How do I run vllm with multiple GPUs using tensor parallelism?"},{"a":"vllm-server documents quantization techniques including AWQ and GPTQ to reduce GPU memory requirements. These methods compress model weights while maintaining inference quality, allowing you to serve larger models or fit more concurrent requests on available VRAM. vllm-server explains how to load pre-quantized models, configure quantization parameters, and measure the trade-offs between memory savings and accuracy for your workload.","q":"What quantization methods does vllm-server support for memory optimization?"},{"a":"vllm-server guides you through configuring OpenAI-compatible API endpoints so your self-hosted models work with existing client libraries and applications. Launch vLLM with the API server enabled, specify your model and port, and vllm-server shows you how to test endpoints using standard OpenAI client code. This compatibility layer lets you swap between cloud providers and self-hosted inference without changing application code.","q":"How do I set up an OpenAI-compatible API endpoint with vllm-server?"},{"a":"vllm-server provides monitoring guidance including Prometheus metrics integration to track throughput, latency, GPU utilization, and memory consumption. It covers troubleshooting common issues like CUDA out-of-memory errors, identifying bottlenecks, and tuning batch size and context window settings. vllm-server helps you correlate metrics with performance problems and adjust configurations\u2014such as reducing max model length or enabling quantization\u2014to resolve resource constraints.","q":"How do I monitor vllm performance and troubleshoot resource issues?"},{"a":"vllm-server focuses on throughput and latency optimization through continuous batching, paged attention, tensor parallelism, and quantization. It explains how to tune batch size, prefill/decode ratios, and GPU memory allocation to balance request concurrency with response time. vllm-server includes benchmarking guidance to measure improvements and configuration recommendations for models like Llama and Mistral serving high-volume production traffic.","q":"Can vllm-server help optimize throughput and latency for high-volume inference?"}],"shadow_tags":["gpu-inference-engine","model-serving-platform","distributed-llm-deployment","openai-api-compatible","memory-efficient-quantization","batch-processing-optimization","multi-gpu-parallelism","production-llm-infrastructure","containerized-deployment","performance-benchmarking"],"summary_rewrite":"vllm-server guides you through deploying and configuring vLLM\u2014a high-performance open-source LLM serving engine\u2014for production workloads. Set up continuous batching, multi-GPU tensor parallelism, model quantization, and OpenAI-compatible API endpoints to serve models like Llama and Mistral at scale. Includes Docker deployment, performance tuning, monitoring with Prometheus metrics, and troubleshooting for common VRAM and throughput issues."},"files":[{"bytes":6223,"path":"infrastructure/local-ai/vllm-server/SKILL.md","sha256":"86f37712761f9c0a8a5dcb713c2a79f031af3941c8a8773a0d6466b6b5799e2a","url":"https://skillfed.io/files/BagelHole/DevOps-Security-Agent-Skills/vllm-server/eb0f7a4e/SKILL.md"}],"id":"BagelHole/DevOps-Security-Agent-Skills/vllm-server","links":{"html":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/vllm-server","md":"https://skillfed.io/BagelHole/DevOps-Security-Agent-Skills/vllm-server.md","repo":"https://github.com/BagelHole/DevOps-Security-Agent-Skills"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":4,"language":"Shell","last_updated":"2026-05-22","license":"MIT","name":"vllm-server","publisher":"BagelHole","stars":44},"relations":{"similar":[{"id":"Orchestra-Research/AI-Research-SKILLs/vllm"},{"id":"OpenLAIR/dr-claw/vllm"},{"id":"moltis-org/moltis/serving-llms-vllm"},{"id":"NousResearch/hermes-agent/serving-llms-vllm"},{"id":"synthetic-sciences/openscience/vllm"},{"id":"graniet/kheish/vllm"},{"id":"ancoleman/ai-design-components/model-serving"},{"id":"agentsope/SkillAlchemy/agentsop-vllm"},{"id":"Prism-Shadow/penguin-harness/vllm"},{"id":"Aradotso/ai-agent-skills/qwen-agentworld-language-world-model"}]},"slug":{"owner":"BagelHole","repo":"DevOps-Security-Agent-Skills","skill":"vllm-server"},"version":"eb0f7a4e"}
