$npx skillfedfor your agent

tpu-inference

With conditionsPyPI Artificial IntelligenceReleased Jul 202679.2K downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — tpu_inference-0.26.0-py3-none-any.whl
v0.26.0 · released 2026-07-31 · Python >=3.10 · 24 runtime deps: tpu-info, yapf, pytest, pytest-mock, absl-py, numpy, google-cloud-storage, jax

Yes, if you have access to Google TPU hardware and need to serve large language models at scale. The package is actively maintained, permissively licensed, and offers low install friction. However, it is strictly tied to TPU infrastructure—it cannot run on CPU or GPU systems. Verify that your target models appear in the tested support matrix and that required features are marked passing rather than experimental or untested.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Google TPU hardware (v3, v4, v5p, v5e, v6e, or v7x) and a compatible TPU environment; cannot run on CPU or GPU systems.
  • Low install friction with a pure-Python wheel distribution.
  • Active maintenance with a recent release 14 days ago.

License · maintenance · safety

permissive license (permissive) — Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.

last release 2026-07-31 (14 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 79,219 downloads/mo, #14,373 on PyPI

Verify before relying

pip install tpu-inference

from tpu_inference import TPUInference
# Requires vLLM and TPU hardware to instantiate and serve models
  • Whether tpu-inference can be installed and imported without TPU hardware present for development/testing purposes.
  • Exact performance characteristics and throughput improvements compared to other TPU serving solutions.
  • Whether all 24 runtime dependencies are strictly required or if some are optional for specific use cases.
Same gist for agents: .md · .json

What it is and what it does

tpu-inference is a vLLM plugin that brings unified inference serving to Google TPUs, bridging PyTorch and JAX model ecosystems under a single lowering path. It allows developers to run PyTorch model definitions natively on TPU hardware without code changes, while also extending native JAX support, all while maintaining vLLM's standard user interface and telemetry. The package targets TPU generations v3 through v7x, with v5e, v6e, and v7x as the recommended targets.

The plugin is designed for production LLM serving workloads. It includes support for core inference features like async scheduling, chunked prefill, KV cache offload, prefix caching, and multimodal inputs. The fact sheet documents tested models including Gemma, Llama, and Qwen families, though some advanced features remain experimental or untested. Installation requires compatible TPU hardware and brings in dependencies across the JAX, PyTorch, and Google Cloud ecosystems.

Use it for

  • Serve open-source LLMs like Llama 3.1/3.3 or Gemma on TPU infrastructure for production inference workloads.
  • Run PyTorch-defined models on TPU hardware without rewriting model code, leveraging TPU performance.
  • Deploy multimodal models (vision-language) on TPUs with tested support for Gemma-4 and Qwen VL variants.
  • Build cost-optimized inference services using TPU's price-to-performance characteristics compared to GPU alternatives.
  • Develop JAX-native inference pipelines with unified backend support alongside PyTorch workloads.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you have access to Google TPU hardware and need to serve large language models at scale.

The package is actively maintained, permissively licensed, and offers low install friction. However, it is strictly tied to TPU infrastructure—it cannot run on CPU or GPU systems. Verify that your target models appear in the tested support matrix and that required features are marked passing rather than experimental or untested.

Install

tpu-inference on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release 14 days ago. Depends on 24 runtime packages including JAX, PyTorch ecosystem libraries, and Google Cloud integrations, which may require significant disk space and compatible system setup.

Requires Google TPU hardware (v3, v4, v5p, v5e, v6e, or v7x) and a compatible TPU environment; cannot run on CPU or GPU systems.

License in practice

Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.

Quickstart

pip install tpu-inference

from tpu_inference import TPUInference
# Requires vLLM and TPU hardware to instantiate and serve models

Verify before relying

  • Whether tpu-inference can be installed and imported without TPU hardware present for development/testing purposes.
  • Exact performance characteristics and throughput improvements compared to other TPU serving solutions.
  • Whether all 24 runtime dependencies are strictly required or if some are optional for specific use cases.

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
24 packages
tpu-infoyapfpytestpytest-mockabsl-pynumpygoogle-cloud-storagejaxjaxliblibtpujaxtypingfastapiflaxtorchaxqwixtorchvisionpathwaysutilsparameterizednumbarunai-model-streamergcsfshypothesissortedcontainerstransformers
MaintenanceActively maintained 14 days since the last release
First released
Downloads79,219 / month, #14,373 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Topic :: Scientific/Engineering :: Artificial Intelligence

Evidence: tpu_inference-0.26.0-py3-none-any.whl

Tags

Capabilities
TPU inference servingvLLM TPU backendlarge language model TPU deploymentJAX PyTorch TPU unifiedTPU LLM inference pluginvLLM hardware acceleration TPUTPU model serving framework
Topics
tpu-servingllm-inferencejax-pytorch-unified

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “TPU inference serving”

  • tpu-inferencetpu-inference is a hardware plugin for vLLM that enables…
  • vllm-tpuvllm-tpu is a high-throughput LLM inference and serving engine…
  • libtpulibtpu is the runtime library that enables JAX, PyTorch, and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also vllm-tpu · libtpu · vllm · tpu-info · gpt-oss · torchax · vllm-cpu · tritonclient · tokamax · llmcompressor