{"enrichment":{"faq":[{"a":"Llm Inference guides you through deploying language models with production-grade inference engines tailored to your hardware and use case. It covers choosing between vLLM for maximum throughput on GPUs, llama.cpp for CPU and edge devices, or Ollama for quick local experimentation, along with quantization strategies and memory optimization.","q":"What is Llm Inference and what can it do?"},{"a":"Llm Inference teaches you to run language model inference by selecting the right engine for your setup\u2014vLLM for GPU throughput, llama.cpp for CPU/edge deployment, or Ollama for local testing. The skill covers configuration, quantization techniques, and platform-specific tuning to balance speed and resource constraints.","q":"How do I run language model inference?"},{"a":"Llm Inference covers three primary engines: vLLM for maximum GPU throughput, llama.cpp for CPU and edge device deployment, and Ollama for quick local experimentation. Each engine is optimized for different hardware configurations and use cases.","q":"What inference engines does Llm Inference support?"},{"a":"Llm Inference teaches optimization through quantization strategies, memory optimization techniques, and platform-specific tuning. These approaches help you balance inference speed with resource constraints, whether running on GPUs, CPUs, or edge devices.","q":"How can I optimize or accelerate language model inference?"},{"a":"Llm Inference guides hardware-specific setup by matching your infrastructure to the right engine: vLLM for GPU systems, llama.cpp for CPU-based or edge environments, or Ollama for local experimentation. Configuration includes quantization and memory optimization tailored to your constraints.","q":"How do I set up Llm Inference for my hardware?"},{"a":"Llm Inference teaches you to build and deploy inference services using production-grade engines. It covers service architecture, engine selection based on hardware, quantization for efficiency, and optimization techniques to ensure fast, resource-efficient language model predictions at scale.","q":"What is an llm inference service and how does it work?"}],"shadow_tags":["model-inference","language-model-execution","ai-prediction","neural-network-inference","text-generation-engine"],"summary_rewrite":"This skill guides you through deploying language models with production-grade inference engines tailored to your hardware and use case. Choose between vLLM for maximum throughput on GPUs, llama.cpp for CPU and edge devices, or Ollama for quick local experimentation. Learn quantization strategies, memory optimization, and platform-specific tuning to balance speed and resource constraints."},"gist":{"api_url":"https://skillfed.io/api/skills/eyadsibai/ltk/llm-inference.json","as_of":"2026-01-15","description":"LLM Inference helps you run language models efficiently across GPUs, CPUs. 6 stars \u00b7 updated Jan 2026. npx skillfed install eyadsibai/ltk/llm-inference","install":{"manual":["git clone https://github.com/eyadsibai/ltk","cp -r ltk ~/.claude/skills/llm-inference"],"primary":"npx skillfed install eyadsibai/ltk/llm-inference","version":"977455f2"},"kind":"skill","mirror_url":"https://skillfed.io/eyadsibai/ltk/llm-inference.md","similar":[{"id":"agentsope/SkillAlchemy/agentsop-llm-engine-selection","name":"agentsop-llm-engine-selection","publisher":"agentsope/SkillAlchemy","url":"https://skillfed.io/agentsope/SkillAlchemy/agentsop-llm-engine-selection"},{"id":"Orchestra-Research/AI-Research-SKILLs/llama-cpp","name":"llama-cpp","publisher":"Orchestra-Research/AI-Research-SKILLs","url":"https://skillfed.io/Orchestra-Research/AI-Research-SKILLs/llama-cpp"},{"id":"graniet/kheish/llama-cpp","name":"Llama Cpp","publisher":"graniet/kheish","url":"https://skillfed.io/graniet/kheish/llama-cpp"}],"title":"Llm Inference by eyadsibai: Run inference on a language \u2014 SkillFed","use":{"when":["Llm Inference teaches you to run language model inference by selecting the right engine for your setup\u2014vLLM for GPU throughput.","Llm Inference covers three primary engines: vLLM for maximum GPU throughput, llama.cpp for CPU and edge device deployment."]},"what":{"lead":"LLM Inference helps you run language models efficiently across GPUs, CPUs, and edge devices using optimized serving engines.","rest":"This skill guides you through deploying language models with production-grade inference engines tailored to your hardware and use case. Choose between vLLM for maximum throughput on GPUs, llama.cpp for CPU and edge devices, or Ollama for quick local experimentation. Learn quantization strategies, memory optimization, and platform-specific tuning to balance speed and resource constraints."}},"id":"eyadsibai/ltk/llm-inference","install":{"mode":"external","repo":"https://github.com/eyadsibai/ltk"},"links":{"html":"https://skillfed.io/eyadsibai/ltk/llm-inference","md":"https://skillfed.io/eyadsibai/ltk/llm-inference.md","repo":"https://github.com/eyadsibai/ltk"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":1,"language":"Python","last_updated":"2026-01-15","license":null,"name":"Llm Inference","publisher":"eyadsibai","stars":6},"relations":{"similar":[{"id":"moltis-org/moltis/llama-cpp"},{"id":"synthetic-sciences/openscience/gguf"},{"id":"OpenLAIR/dr-claw/gguf"},{"id":"Orchestra-Research/AI-Research-SKILLs/gguf"},{"id":"graniet/kheish/gguf"},{"id":"agentsope/SkillAlchemy/agentsop-llm-engine-selection"},{"id":"synthetic-sciences/openscience/llama-cpp"},{"id":"Orchestra-Research/AI-Research-SKILLs/llama-cpp"},{"id":"OpenLAIR/dr-claw/llama-cpp"},{"id":"graniet/kheish/llama-cpp"}]},"slug":{"owner":"eyadsibai","repo":"ltk","skill":"llm-inference"},"version":"977455f2"}
