$npx skillfedfor your agent

mooncake-transfer-engine-cuda13

A KVCache-centric Disaggregated Architecture for large-scale LLM inference and training. (CUDA 13 version)

With conditionsPyPI Artificial IntelligenceReleased Jul 2026158.0K downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — mooncake_transfer_engine_cuda13-0.3.12.post1-cp310-cp310-manylinux_2_28_aarch64.whl · mooncake_transfer_engine_cuda13-0.3.12.post1-cp310-cp310-manylinux_2_28_x86_64.whl · mooncake_transfer_engine_cuda13-0.3.12.post1-cp311-cp311-manylinux_2_28_aarch64.whl
v0.3.12.post1 · released 2026-07-25 · Python >=3.10 · 3 runtime deps: aiohttp, requests, msgpack

Yes, if you are deploying large-scale LLM inference or training on multi-node GPU clusters and need efficient distributed KV cache management. The package is production-stable, actively maintained, and integrates with established frameworks (vLLM, SGLang, TensorRT LLM). Requires CUDA 13 and Python 3.10+; not suitable for single-node or CPU-only setups.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires CUDA 13 runtime, NVIDIA GPU with RDMA support, and manylinux_2_28 compatible Linux system (x86_64 or aarch64).
  • Python 3.10 or later.
  • Medium install friction due to CUDA 13 requirement and platform-specific wheels (manylinux_2_28 for x86_64 and aarch64).

License · maintenance · safety

permissive license (permissive) — Permissive license treatment allows commercial and private use without significant restrictions.

last release 2026-07-25 (20 days) · last repo commit 2026-08-14 · 6,276 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 158,000 downloads/mo, #10,738 on PyPI

Verify before relying

pip install mooncake-transfer-engine-cuda13

import mooncake_transfer_engine_cuda13
# Use with vLLM or SGLang via their Mooncake connectors
# or directly via the Transfer Engine API for custom RDMA transfers
  • Whether CUDA 13 runtime is required on the system or if the package bundles it
  • Specific performance gains or throughput numbers for typical workloads
  • Compatibility with non-NVIDIA GPU architectures or inference frameworks beyond those mentioned
Same gist for agents: .md · .json

What it is and what it does

Mooncake Transfer Engine is a component of the Mooncake infrastructure platform designed for large-scale LLM inference and training. It implements a KV cache-centric disaggregated architecture that separates prefill and decode compute clusters while using underutilized CPU, DRAM, and SSD resources to build a distributed KV cache pool. The engine provides zero-copy RDMA-based transfer of KV cache data across GPU clusters, enabling efficient cross-instance sharing of cached key-value states.

The package is tightly integrated with major LLM serving frameworks including vLLM, SGLang, TensorRT LLM, and LMDeploy. It serves as a backend for distributed KV cache management, hierarchical caching across device/host/remote storage tiers, and disaggregated prefill-decode inference patterns. The transfer engine handles large-scale model updates and multimodal embedding distribution in production training and inference pipelines.

Use it for

  • Disaggregated LLM inference on multi-node GPU clusters to separate compute-intensive prefill from decode operations
  • Distributed KV cache pooling to reduce per-node memory pressure and enable higher request throughput
  • Zero-copy cross-GPU weight and embedding transfer during large-scale distributed training
  • Hierarchical KV cache storage with intelligent offloading across device, host, and remote tiers
  • Multimodal inference with efficient cross-instance sharing of encoder outputs (e.g., Vision Transformer embeddings)

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are deploying large-scale LLM inference or training on multi-node GPU clusters and need efficient distributed KV cache management.

The package is production-stable, actively maintained, and integrates with established frameworks (vLLM, SGLang, TensorRT LLM). Requires CUDA 13 and Python 3.10+; not suitable for single-node or CPU-only setups.

Install

mooncake-transfer-engine-cuda13 on PyPI

Before you install

Medium install friction due to CUDA 13 requirement and platform-specific wheels (manylinux_2_28 for x86_64 and aarch64). Package is actively maintained with recent commits and production-stable status. Requires Python 3.10+.

Requires CUDA 13 runtime, NVIDIA GPU with RDMA support, and manylinux_2_28 compatible Linux system (x86_64 or aarch64). Python 3.10 or later.

License in practice

Permissive license treatment allows commercial and private use without significant restrictions.

Quickstart

pip install mooncake-transfer-engine-cuda13

import mooncake_transfer_engine_cuda13
# Use with vLLM or SGLang via their Mooncake connectors
# or directly via the Transfer Engine API for custom RDMA transfers

Verify before relying

  • Whether CUDA 13 runtime is required on the system or if the package bundles it
  • Specific performance gains or throughput numbers for typical workloads
  • Compatibility with non-NVIDIA GPU architectures or inference frameworks beyond those mentioned

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
3 packages
aiohttprequestsmsgpack
MaintenanceActively maintained 20 days since the last release
Last repo commit
First released
Downloads158,000 / month, #10,738 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: GPU :: NVIDIA CUDAIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: Implementation :: CPythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: System :: Distributed Computing

Evidence: mooncake_transfer_engine_cuda13-0.3.12.post1-cp310-cp310-manylinux_2_28_aarch64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp310-cp310-manylinux_2_28_x86_64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp311-cp311-manylinux_2_28_aarch64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp311-cp311-manylinux_2_28_x86_64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_aarch64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp313-cp313-manylinux_2_28_aarch64.whl; mooncake_transfer_engine_cuda13-0.3.12.post1-cp313-cp313-manylinux_2_28_x86_64.whl

Tags

Capabilities
llm kv cache transferrdma distributed inferencedisaggregated llm servinggpu cluster data transferkv cache pool management
Topics
llm-inferencedistributed-systemsgpu-acceleration
PyPI keywords
mooncaketransfer enginekv cachellm inferencerdmacuda13

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “llm kv cache transfer”

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also mooncake-transfer-engine · memcache-hybrid · lmcache · vllm · vllm-router · nvidia-nvshmem-cu13 · vllm-tpu · vllm-cpu · vineyard · vineyard-bdist

Further reading