{"enrichment":{"faq":[{"a":"agentsop-llm-engine-selection guides you through a five-step decision workflow that analyzes your hardware topology, workload profile, and operational constraints. The skill maps seven engines\u2014vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX\u2014to their strengths across GPU clusters, edge devices, and single-user scenarios, helping you eliminate incompatible options and identify your best candidates before benchmarking.","q":"Which LLM inference engine should I use for my specific hardware and workload?"},{"a":"agentsop-llm-engine-selection compares these engines across performance tradeoffs and benchmarking methodology. vLLM excels in multi-user GPU cluster deployments with high throughput; SGLang adds structured output and agent-friendly features; TensorRT-LLM offers maximum performance on NVIDIA hardware but requires more optimization expertise. The skill helps you weigh these differences against your specific constraints\u2014latency, throughput, hardware topology, and operational overhead.","q":"How do vLLM, SGLang, and TensorRT-LLM compare for production LLM serving?"},{"a":"agentsop-llm-engine-selection evaluates llama.cpp, Ollama, and MLX for resource-constrained environments. llama.cpp and Ollama suit CPU-only servers with minimal dependencies; MLX is optimized for Apple Silicon. The skill's decision workflow helps you match engine capabilities\u2014quantization support, model format compatibility, and inference speed\u2014to your edge hardware and latency requirements.","q":"What inference engine should I choose for CPU-only or edge deployment?"},{"a":"agentsop-llm-engine-selection provides a migration evaluation framework by comparing runtime performance, feature gaps, and operational switching costs. The skill assesses whether your current bottleneck is throughput, latency, structured output support, or cost, then determines if migration gains justify redeployment effort. It guides you through benchmarking your top candidates on your actual workload before committing to a switch.","q":"Should I migrate from TGI to vLLM or another serving runtime?"},{"a":"agentsop-llm-engine-selection supports heterogeneous deployment design by mapping each engine to its optimal tier\u2014vLLM or TensorRT-LLM for GPU clusters, MLX for Apple Silicon, llama.cpp or Ollama for CPU fallback. The skill helps you route requests intelligently, balance load across tiers, and select engines that share compatible model formats to minimize conversion overhead and operational complexity.","q":"How do I design a multi-tier LLM serving deployment across heterogeneous hardware?"},{"a":"agentsop-llm-engine-selection explains performance tradeoff evaluation across inference engines, emphasizing that benchmark headlines often hide hardware-specific tuning and workload assumptions. The skill guides you to benchmark your top candidates on your actual hardware, model size, batch profile, and latency requirements rather than relying on published results, ensuring fair comparison and production-realistic performance estimates.","q":"What is the right LLM inference benchmarking methodology for engine comparison?"}],"shadow_tags":["inference-runtime-selection","hardware-workload-matching","llm-deployment-architecture","serving-stack-comparison","production-inference-patterns","gpu-topology-constraints","structured-generation-routing","edge-inference-options","multi-tier-serving-design","vendor-lock-in-analysis"],"summary_rewrite":"This skill guides you through selecting an LLM serving engine by analyzing hardware topology, workload profile, and operational constraints rather than benchmark headlines. It maps seven engines\u2014vLLM, SGLang, TensorRT-LLM, TGI, llama.cpp, Ollama, and MLX\u2014to their strengths across GPU clusters, edge devices, and single-user scenarios, then walks you through a five-step decision workflow to eliminate incompatible options and benchmark your top candidates."},"files":[{"bytes":24815,"path":"skills/agentsop-llm-engine-selection/SKILL.md","sha256":"4b7866571be9738eac5dc34d823cdc3d354c3b5eff39b9b5952712975cc05185","url":"https://skillfed.io/files/agentsope/SkillAlchemy/agentsop-llm-engine-selection/bfcb837c/SKILL.md"}],"id":"agentsope/SkillAlchemy/agentsop-llm-engine-selection","links":{"html":"https://skillfed.io/agentsope/SkillAlchemy/agentsop-llm-engine-selection","md":"https://skillfed.io/agentsope/SkillAlchemy/agentsop-llm-engine-selection.md","repo":"https://github.com/agentsope/SkillAlchemy"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":12,"language":"Python","last_updated":"2026-06-30","license":"MIT","name":"agentsop-llm-engine-selection","publisher":"agentsope","stars":219},"relations":{"similar":[{"id":"agentsope/SkillAlchemy/agentsop-vllm"},{"id":"eyadsibai/ltk/llm-inference"},{"id":"vasilyu1983/AI-Agents-public/ai-llm-inference"},{"id":"synthetic-sciences/openscience/tensorrt-llm"},{"id":"NousResearch/hermes-agent/tensorrt-llm"},{"id":"Orchestra-Research/AI-Research-SKILLs/tensorrt-llm"},{"id":"OpenLAIR/dr-claw/tensorrt-llm"},{"id":"agentsope/SkillAlchemy/agentsop-framework-selection"},{"id":"agentsope/SkillAlchemy/agentsop-selfhost-decision"},{"id":"synthetic-sciences/openscience/vllm"}]},"slug":{"owner":"agentsope","repo":"SkillAlchemy","skill":"agentsop-llm-engine-selection"},"version":"bfcb837c"}
