lmcache
A LLM serving engine extension to reduce TTFT and increase throughput, especially under long-context scenarios.
Decision gist · record as of 2026-08-14
Yes, with conditions. Install if you run LLM inference at scale with multi-turn, RAG, or long-context workloads and want to reduce latency and cost. Active maintenance, permissive license, and strong community adoption are positive signals. Medium install friction (37 dependencies, Linux-only wheels) and alpha status mean you should test integration with your serving engine and storage backend before production deployment. No known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Linux x86_64, Python 3.10–3.13, and torch; GPU recommended for typical LLM inference workloads.
- Medium install friction: 37 runtime dependencies including torch, numpy, fastapi, redis, and observability libraries (opentelemetry, prometheus).
- Wheels available for Python 3.10–3.13 on Linux x86_64.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license and copyright notices in distributions.
last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 11,147 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 83,225 downloads/mo, #14,090 on PyPI
Alternatives
Verify before relying
pip install lmcache
from lmcache import LMCache
# Initialize cache manager
cache = LMCache()- Specific integration steps with mainstream serving engines beyond the high-level architecture.
- Performance gains (TTFT reduction, throughput improvement) quantified for representative workloads.
- Compatibility and tested versions for each supported storage backend (Redis, S3, local disk, etc.).
- Overhead of the daemon process and observability stack on latency-sensitive inference.
What it is and what it does
LMCache decouples KV cache management from inference engines, storing transformer caches persistently across GPU, CPU, and remote storage tiers. It enables reuse of cached key-value blocks across requests and sessions, reducing redundant prefill computation for multi-turn conversations, agentic workloads, and retrieval-augmented generation (RAG). The layer runs as a standalone daemon, survives engine restarts, and provides observability through Kubernetes metrics, cache hit rates, and token-level diagnostics.
The package integrates with mainstream open-source serving engines and supports pluggable storage backends (CPU RAM, local SSD, Redis, S3-compatible systems, and others) and transport layers (NVLink, RDMA, TCP). It includes non-prefix KV reuse via CacheBlend for selective recomputation, prefill-decode disaggregation, and a flexible serialization interface for compression and token dropping. Vendor neutrality means you can switch storage or serving engines without losing cached KV data.
Use it for
- Multi-turn conversational AI: cache KV blocks from earlier turns to skip prefill recomputation on follow-up messages.
- RAG and knowledge-augmented generation: reuse cached context embeddings across queries with the same knowledge base.
- Agentic workloads: share KV caches across multiple reasoning steps and tool calls within a single session.
- Long-context inference: offload KV caches to CPU or remote storage when GPU memory is exhausted, enabling longer prompts.
- Batch serving with cache sharing: reduce total prefill cost across requests that share common prompt prefixes.
- Cross-engine failover: persist KV caches so inference can resume on a different serving engine without recomputation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Install if you run LLM inference at scale with multi-turn, RAG, or long-context workloads and want to reduce latency and cost. Active maintenance, permissive license, and strong community adoption are positive signals. Medium install friction (37 dependencies, Linux-only wheels) and alpha status mean you should test integration with your serving engine and storage backend before production deployment. No known vulnerabilities.
Install
lmcache on PyPI
Before you install
Medium install friction: 37 runtime dependencies including torch, numpy, fastapi, redis, and observability libraries (opentelemetry, prometheus). Wheels available for Python 3.10–3.13 on Linux x86_64. Active maintenance with recent release (9 days old) and strong community signal.
Requires Linux x86_64, Python 3.10–3.13, and torch; GPU recommended for typical LLM inference workloads.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license and copyright notices in distributions.
Quickstart
pip install lmcache
from lmcache import LMCache
# Initialize cache manager
cache = LMCache()
Verify before relying
- Specific integration steps with mainstream serving engines beyond the high-level architecture.
- Performance gains (TTFT reduction, throughput improvement) quantified for representative workloads.
- Compatibility and tested versions for each supported storage backend (Redis, S3, local disk, etc.).
- Overhead of the daemon process and observability stack on latency-sensitive inference.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release <3.14,>=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 37 packagesaiofileaiofilesblake3aiohttpawscrtcryptographycufile-pythonfastapihttpxhuggingface_hubmsgspecnumpynumbanvtxopentelemetry-apiopentelemetry-sdkopentelemetry-exporter-otlpopentelemetry-exporter-prometheusprometheus_clientpsutilpy-cpuinfopytestpyyamlpyzmqredissafetensorssetuptoolssetuptools_scmsortedcontainerstorch |
| Maintenance | Actively maintained 9 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 83,225 / month, #14,090 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaEnvironment :: GPUIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: Implementation :: CPythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information Analysis |
Evidence: lmcache-0.5.3-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “reduce ttft large language models”
- lmcacheLMCache is a KV cache management layer that stores, reuses, and…
- loralibloralib provides PyTorch modules that implement Low-Rank Adaptation…
- peftPEFT implements parameter-efficient fine-tuning methods (LoRA, QLoRA,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also mooncake-transfer-engine · mooncake-transfer-engine-cuda13 · vllm · memcache-hybrid · vllm-tpu · gptcache · vllm-cpu · ipex-llm · llmcompressor · vllm-router