--- id: eyadsibai/ltk/llm-inference version: "977455f2" license: none install: manual updated: 2026-01-15 --- # Llm Inference — This skill guides you through deploying language models with production-grade inference engines tailored to your hardware and use case. Choose between vLLM for maximum throughput on GPUs, llama.cpp for CPU and edge devices, or Ollama for quick local experimentation. Learn quantization strategies, memory optimization, and platform-specific tuning to balance speed and resource constraints. Publisher: eyadsibai · Stars: 6 · Updated: 2026-01-15 Install (manual): `git clone https://github.com/eyadsibai/ltk` *Unlicensed repository - metadata only, no file contents reproduced.* [View on SkillFed](https://skillfed.io/eyadsibai/ltk/llm-inference) · [View on GitHub](https://github.com/eyadsibai/ltk)