instanttensor
An ultra-fast, distributed Safetensors loader
Decision gist · record as of 2026-08-14
Yes, if you load large Safetensors models onto GPU and have either high storage bandwidth, constrained host memory, or frequent model switching. The package is actively maintained, permissively licensed, and already integrated into vLLM. Install friction is moderate (compiled wheels, torch dependency). Not recommended if you load small models infrequently or have ample host memory for caching—standard Safetensors loading will suffice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA or ROCm GPU and torch installed; Linux x86_64 only in current wheels.
- Medium friction: compiled wheels available for Python 3.10–3.14 on Linux x86_64, but requires torch as a runtime dependency.
- Active maintenance status with recent release.
License · maintenance · safety
permissive license (permissive) — Apache License 2.0 (permissive): you may use, modify, and distribute freely provided you include a copy of the license and document any changes. No restrictions on commercial use.
last release 2026-05-27 (79 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 92,667 downloads/mo, #13,435 on PyPI
Alternatives
Verify before relying
pip install instanttensor
from instanttensor import safe_open
with safe_open("model.safetensors", framework="pt", device=0) as f:
for name, tensor in f.tensors():
print(name, tensor.shape)- Whether zero-copy mode (copy=False) is production-ready or still experimental given alpha status.
- Performance gains on non-H200/H100 hardware or with smaller models.
- Stability of distributed loading with subgroups across different parallelism strategies.
What it is and what it does
InstantTensor is a Safetensors loader built to maximize I/O throughput when loading model weights onto GPU. It uses direct I/O, tuned concurrency, and pipelining to avoid slow page cache allocation, and supports distributed loading via torch.distributed NCCL for coordinated multi-GPU reads. The package is designed for scenarios where models are large, storage bandwidth is high, or the model cannot be cached in host memory—such as when memory is consumed by KV cache offloading in LLM serving, or when loading multiple models that cannot fit simultaneously.
The library exposes a `safe_open` context manager that yields tensors from Safetensors files, with options for zero-copy streaming into preallocated buffers, backend selection (AIO, URING, CUFILE, MMAP), and buffered vs. direct I/O modes. It integrates with torch and requires a GPU platform (CUDA or ROCm). The package is in alpha status and depends only on torch.
Use it for
- Loading large language models (30B+) onto single or multi-GPU setups where cold-start latency matters.
- Serving scenarios where host memory is constrained by KV cache offloading or other allocations.
- Multi-model serving where models are switched frequently and cannot be cached together.
- Distributed inference with tensor parallelism (TP=8+) where each GPU receives small, non-contiguous shards.
- Loading model checkpoints from tmpfs or high-bandwidth storage (≥5 GB/s) where direct I/O is beneficial.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you load large Safetensors models onto GPU and have either high storage bandwidth, constrained host memory, or frequent model switching.
The package is actively maintained, permissively licensed, and already integrated into vLLM. Install friction is moderate (compiled wheels, torch dependency). Not recommended if you load small models infrequently or have ample host memory for caching—standard Safetensors loading will suffice.
Install
instanttensor on PyPI
Before you install
Medium friction: compiled wheels available for Python 3.10–3.14 on Linux x86_64, but requires torch as a runtime dependency. Active maintenance status with recent release.
Requires CUDA or ROCm GPU and torch installed; Linux x86_64 only in current wheels.
License in practice
Apache License 2.0 (permissive): you may use, modify, and distribute freely provided you include a copy of the license and document any changes. No restrictions on commercial use.
Quickstart
pip install instanttensor
from instanttensor import safe_open
with safe_open("model.safetensors", framework="pt", device=0) as f:
for name, tensor in f.tensors():
print(name, tensor.shape)
Verify before relying
- Whether zero-copy mode (copy=False) is production-ready or still experimental given alpha status.
- Performance gains on non-H200/H100 hardware or with smaller models.
- Stability of distributed loading with subgroups across different parallelism strategies.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packagetorch |
| Maintenance | Actively maintained 79 days since the last release |
| First released | |
| Downloads | 92,667 / month, #13,435 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchProgramming Language :: C++Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: instanttensor-0.1.9-cp310-cp310-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; instanttensor-0.1.9-cp311-cp311-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; instanttensor-0.1.9-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; instanttensor-0.1.9-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; instanttensor-0.1.9-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “fast model weight loading”
- instanttensorInstantTensor provides a high-throughput Safetensors loader optimized…
- auto-roundAutoRound quantizes large language models and vision-language models…
- transformer-lensTransformerLens loads and inspects the internal activations of…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also fastsafetensors · safetensors · tensorizer · comfy-aimdo · torch · spmd-types · compressed-tensors · nvidia-nccl-cu13 · nvidia-cufile · nvidia-cufile-cu12