tensorizer
A tool for fast PyTorch module, model, and tensor serialization + deserialization.
Decision gist · record as of 2026-08-14
Yes, if you deploy large PyTorch models in containerized or serverless environments and need to minimize load latency. The permissive MIT license, active maintenance, and zero known vulnerabilities make it safe to adopt. Install friction is low. Not necessary for single-machine development or models small enough to embed in container images.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- libnacl requires system libsodium library; boto3 credentials needed for S3 access; torch and CUDA/CPU device setup required for deserialization.
- Low friction install with a pure-Python wheel.
- Maintenance is active with recent commits and stable production status.
License · maintenance · safety
MIT License (permissive) — MIT License (permissive) allows commercial and private use with minimal restrictions; attribution required but no copyleft obligations.
last release 2026-04-21 (115 days) · last repo commit 2026-07-07 · 321 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 160,437 downloads/mo, #10,661 on PyPI
Alternatives
Verify before relying
pip install tensorizer
from tensorizer import TensorDeserializer
import torch
deserializer = TensorDeserializer("s3://bucket/model.tensors", device="cuda")
deserializer.load_into_module(model)
deserializer.close()- Whether the ~5GB/s wire-speed claim applies to typical network conditions outside 40GbE lab environments.
- Redis support maturity and whether preliminary status affects production reliability.
- Memory overhead during serialization/deserialization of very large models.
What it is and what it does
tensorizer is a PyTorch serialization library designed to decouple large models from container images and enable fast streaming loads from remote storage. It writes torch.nn.Module instances to a binary format that can be read back via HTTP/HTTPS, S3, Redis, or local filesystem, with streaming deserialization that avoids downloading entire models to disk first.
The package is built for serverless and containerized inference workflows where model load time directly impacts latency. By storing serialized models separately from container images, you can update models without rebuilding containers and stream weights directly into GPU memory. It supports local filesystem access for development, S3 for cloud object storage, and HTTP/HTTPS endpoints for any compatible server. Redis support is available but marked preliminary and not recommended for production model deployment.
Use it for
- Reduce KNative serverless function cold-start latency by streaming large models from S3 instead of embedding them in container images.
- Deploy and update multi-gigabyte language models without rebuilding or redeploying container images.
- Load models directly into GPU memory from HTTP/HTTPS endpoints at network wire-speed without intermediate disk storage.
- Share model state between inference pods via Redis for per-request data loading.
- Serialize and deserialize models locally for development and testing with the same fast streaming path used in production.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you deploy large PyTorch models in containerized or serverless environments and need to minimize load latency.
The permissive MIT license, active maintenance, and zero known vulnerabilities make it safe to adopt. Install friction is low. Not necessary for single-machine development or models small enough to embed in container images.
Install
tensorizer on PyPI
Before you install
Low friction install with a pure-Python wheel. Maintenance is active with recent commits and stable production status. Eight runtime dependencies include torch, numpy, and cloud/storage libraries (boto3, redis, hiredis); libnacl requires a system libsodium library.
libnacl requires system libsodium library; boto3 credentials needed for S3 access; torch and CUDA/CPU device setup required for deserialization.
License in practice
MIT License (permissive) allows commercial and private use with minimal restrictions; attribution required but no copyleft obligations.
Quickstart
pip install tensorizer
from tensorizer import TensorDeserializer
import torch
deserializer = TensorDeserializer("s3://bucket/model.tensors", device="cuda")
deserializer.load_into_module(model)
deserializer.close()
Verify before relying
- Whether the ~5GB/s wire-speed claim applies to typical network conditions outside 40GbE lab environments.
- Redis support maturity and whether preliminary status affects production reliability.
- Memory overhead during serialization/deserialization of very large models.
Package facts
| License | MIT License permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagestorchnumpyprotobufpsutilboto3redishiredislibnacl |
| Maintenance | Actively maintained 115 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 160,437 / month, #10,661 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: InternetTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: System :: Distributed Computing |
Evidence: tensorizer-2.12.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pytorch model serialization”
- tensorizerSerializes and deserializes PyTorch modules and tensors to/from HTTP,…
- optimum-quantoA PyTorch quantization backend that reduces model size and memory by…
- torch-emaComputes exponential moving averages of PyTorch model parameters…
Give your agent the search over MCP, or paste the wish link into any chat.
More Internet packages
Botocore provides low-level, data-driven access to Amazon Web Services APIs, serving as the foundation for the AWS CLI and boto3 libraries.
Install it if you need programmatic access to AWS services.
Provides an async client for AWS services using botocore and aiohttp, allowing you to call AWS APIs asynchronously within asyncio-based applications.
Install it if you need to call AWS services from async Python code; it is the standard way to do so.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides a platform-independent file locking mechanism to coordinate access to files across processes and threads.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides common Protocol Buffer message definitions used across Google Cloud APIs, enabling Python clients to interact with Google services.
See also safetensors · torch · instanttensor · comfy-aimdo · s3torchconnector · s3torchconnectorclient · transformers · tensorly · litdata · fastsafetensors