$npx skillfedfor your agent

lmcache

A LLM serving engine extension to reduce TTFT and increase throughput, especially under long-context scenarios.

With conditionsPyPI Artificial IntelligenceReleased Aug 202683.2K downloads / moApache-2.0Platform wheel

Decision gist · record as of 2026-08-14

platform wheels — lmcache-0.5.3-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl · lmcache-0.5.3-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl · lmcache-0.5.3-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
v0.5.3 · released 2026-08-05 · Python <3.14,>=3.10 · 37 runtime deps: aiofile, aiofiles, blake3, aiohttp, awscrt, cryptography, cufile-python, fastapi

Yes, with conditions. Install if you run LLM inference at scale with multi-turn, RAG, or long-context workloads and want to reduce latency and cost. Active maintenance, permissive license, and strong community adoption are positive signals. Medium install friction (37 dependencies, Linux-only wheels) and alpha status mean you should test integration with your serving engine and storage backend before production deployment. No known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Linux x86_64, Python 3.10–3.13, and torch; GPU recommended for typical LLM inference workloads.
  • Medium install friction: 37 runtime dependencies including torch, numpy, fastapi, redis, and observability libraries (opentelemetry, prometheus).
  • Wheels available for Python 3.10–3.13 on Linux x86_64.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license and copyright notices in distributions.

last release 2026-08-05 (9 days) · last repo commit 2026-08-14 · 11,147 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 83,225 downloads/mo, #14,090 on PyPI

Verify before relying

pip install lmcache

from lmcache import LMCache

# Initialize cache manager
cache = LMCache()
  • Specific integration steps with mainstream serving engines beyond the high-level architecture.
  • Performance gains (TTFT reduction, throughput improvement) quantified for representative workloads.
  • Compatibility and tested versions for each supported storage backend (Redis, S3, local disk, etc.).
  • Overhead of the daemon process and observability stack on latency-sensitive inference.
Same gist for agents: .md · .json

What it is and what it does

LMCache decouples KV cache management from inference engines, storing transformer caches persistently across GPU, CPU, and remote storage tiers. It enables reuse of cached key-value blocks across requests and sessions, reducing redundant prefill computation for multi-turn conversations, agentic workloads, and retrieval-augmented generation (RAG). The layer runs as a standalone daemon, survives engine restarts, and provides observability through Kubernetes metrics, cache hit rates, and token-level diagnostics.

The package integrates with mainstream open-source serving engines and supports pluggable storage backends (CPU RAM, local SSD, Redis, S3-compatible systems, and others) and transport layers (NVLink, RDMA, TCP). It includes non-prefix KV reuse via CacheBlend for selective recomputation, prefill-decode disaggregation, and a flexible serialization interface for compression and token dropping. Vendor neutrality means you can switch storage or serving engines without losing cached KV data.

Use it for

  • Multi-turn conversational AI: cache KV blocks from earlier turns to skip prefill recomputation on follow-up messages.
  • RAG and knowledge-augmented generation: reuse cached context embeddings across queries with the same knowledge base.
  • Agentic workloads: share KV caches across multiple reasoning steps and tool calls within a single session.
  • Long-context inference: offload KV caches to CPU or remote storage when GPU memory is exhausted, enabling longer prompts.
  • Batch serving with cache sharing: reduce total prefill cost across requests that share common prompt prefixes.
  • Cross-engine failover: persist KV caches so inference can resume on a different serving engine without recomputation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

Install if you run LLM inference at scale with multi-turn, RAG, or long-context workloads and want to reduce latency and cost. Active maintenance, permissive license, and strong community adoption are positive signals. Medium install friction (37 dependencies, Linux-only wheels) and alpha status mean you should test integration with your serving engine and storage backend before production deployment. No known vulnerabilities.

Install

lmcache on PyPI

Before you install

Medium install friction: 37 runtime dependencies including torch, numpy, fastapi, redis, and observability libraries (opentelemetry, prometheus). Wheels available for Python 3.10–3.13 on Linux x86_64. Active maintenance with recent release (9 days old) and strong community signal.

Requires Linux x86_64, Python 3.10–3.13, and torch; GPU recommended for typical LLM inference workloads.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license and copyright notices in distributions.

Quickstart

pip install lmcache

from lmcache import LMCache

# Initialize cache manager
cache = LMCache()

Verify before relying

  • Specific integration steps with mainstream serving engines beyond the high-level architecture.
  • Performance gains (TTFT reduction, throughput improvement) quantified for representative workloads.
  • Compatibility and tested versions for each supported storage backend (Redis, S3, local disk, etc.).
  • Overhead of the daemon process and observability stack on latency-sensitive inference.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release <3.14,>=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
37 packages
aiofileaiofilesblake3aiohttpawscrtcryptographycufile-pythonfastapihttpxhuggingface_hubmsgspecnumpynumbanvtxopentelemetry-apiopentelemetry-sdkopentelemetry-exporter-otlpopentelemetry-exporter-prometheusprometheus_clientpsutilpy-cpuinfopytestpyyamlpyzmqredissafetensorssetuptoolssetuptools_scmsortedcontainerstorch
MaintenanceActively maintained 9 days since the last release
Last repo commit
First released
Downloads83,225 / month, #14,090 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaEnvironment :: GPUIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: Implementation :: CPythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information Analysis

Evidence: lmcache-0.5.3-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; lmcache-0.5.3-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Tags

Capabilities
kv cache management llm inferencereduce ttft large language modelscache reuse transformer servingllm inference optimizationpersistent kv cache storagemulti-turn conversation cachingrag knowledge augmented generation
Topics
llm-inferencecache-optimizationdistributed-systems

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “reduce ttft large language models”

  • lmcacheLMCache is a KV cache management layer that stores, reuses, and…
  • loralibloralib provides PyTorch modules that implement Low-Rank Adaptation…
  • peftPEFT implements parameter-efficient fine-tuning methods (LoRA, QLoRA,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also mooncake-transfer-engine · mooncake-transfer-engine-cuda13 · vllm · memcache-hybrid · vllm-tpu · gptcache · vllm-cpu · ipex-llm · llmcompressor · vllm-router

Further reading