s3torchconnector
S3 connector integration for PyTorch
Decision gist · record as of 2026-08-14
Yes, if you train PyTorch models on AWS and need to load data from or checkpoint to S3. High install friction (compiled deps, platform-specific wheels) and AWS credential configuration requirements are real costs, but the package is actively maintained, permissively licensed, and eliminates custom S3 integration code. Not worth installing if you're not on Linux/macOS or don't use S3 for training data.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- AWS credentials must be configured (EC2 role, AWS CLI, ~/.aws/credentials, or AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY environment variables).
- Pre-built wheels available only for Linux and macOS; other platforms require building from source.
- High install friction due to compiled dependencies (s3torchconnectorclient) and platform-specific wheels; pre-built wheels available only for Linux and macOS.
License · maintenance · safety
permissive license (permissive) — Permissive license treatment allows commercial and private use without restriction.
last release 2026-02-20 (175 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,355,211 downloads/mo, #4,008 on PyPI
Alternatives
Verify before relying
pip install s3torchconnector
from s3torchconnector import S3IterableDataset
dataset = S3IterableDataset.from_prefix(
"s3://my-bucket/data",
region="us-east-1"
)
for item in dataset:
print(item.key, len(item.read()))- Whether macOS x86_64 wheel support deprecation affects your deployment target
- Python 3.8 support timeline and deprecation schedule for your long-term maintenance
- Performance gains over manual S3 listing and concurrent request management in your workload
What it is and what it does
s3torchconnector bridges PyTorch's data loading pipeline with Amazon S3, eliminating the need to write custom code for bucket listing and concurrent request handling. It implements PyTorch's dataset primitives—both map-style for random access and iterable-style for streaming—so training jobs can fetch data directly from S3 with automatic performance optimization. It also provides checkpoint interfaces to save and load model state directly to S3 without intermediate local storage, and includes distributed checkpoint support via StorageWriter and StorageReader implementations compatible with PyTorch's distributed checkpoint framework.
The package depends on torch and s3torchconnectorclient (a compiled native component), which creates platform-specific installation requirements. It supports Python 3.8 through 3.14 and requires PyTorch 2.0 or newer (PyTorch 2.3+ for distributed checkpoint features). AWS credentials must be configured via EC2 instance role, AWS CLI, credential files, or environment variables before use.
Use it for
- Stream training data from S3 buckets into PyTorch DataLoaders without downloading to local disk first
- Save model checkpoints directly to S3 during distributed training without staging through local storage
- Load distributed checkpoints from S3 using optimized readers for faster recovery in multi-node training
- Access S3 Express One Zone directory buckets for high-performance data access in time-sensitive training jobs
- Implement random-access training datasets backed by S3 for flexible sampling patterns
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you train PyTorch models on AWS and need to load data from or checkpoint to S3.
High install friction (compiled deps, platform-specific wheels) and AWS credential configuration requirements are real costs, but the package is actively maintained, permissively licensed, and eliminates custom S3 integration code. Not worth installing if you're not on Linux/macOS or don't use S3 for training data.
Install
s3torchconnector on PyPI
Before you install
High install friction due to compiled dependencies (s3torchconnectorclient) and platform-specific wheels; pre-built wheels available only for Linux and macOS. Maintenance status is active with a recent release.
AWS credentials must be configured (EC2 role, AWS CLI, ~/.aws/credentials, or AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY environment variables). Pre-built wheels available only for Linux and macOS; other platforms require building from source.
License in practice
Permissive license treatment allows commercial and private use without restriction.
Quickstart
pip install s3torchconnector
from s3torchconnector import S3IterableDataset
dataset = S3IterableDataset.from_prefix(
"s3://my-bucket/data",
region="us-east-1"
)
for item in dataset:
print(item.key, len(item.read()))
Verify before relying
- Whether macOS x86_64 wheel support deprecation affects your deployment target
- Python 3.8 support timeline and deprecation schedule for your long-term maintenance
- Performance gains over manual S3 listing and concurrent request management in your workload
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release <3.15,>=3.8 |
| Install friction | High. Source build required |
| Runtime dependencies | 2 packagestorchs3torchconnectorclient |
| Maintenance | Actively maintained 175 days since the last release |
| First released | |
| Downloads | 1,355,211 / month, #4,008 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableLicense :: OSI Approved :: BSD LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Utilities |
Evidence: s3torchconnector-1.5.0.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pytorch s3 data loading”
- s3torchconnectorProvides PyTorch dataset primitives and checkpoint interfaces for…
- s3torchconnectorclientInternal S3 client implementation providing optimized data loading…
- tensorizerSerializes and deserializes PyTorch modules and tensors to/from HTTP,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Utilities packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Pygments is a syntax highlighter that colorizes source code and text in over 500 languages and formats, outputting to HTML, LaTeX, RTF, SVG, images, or ANSI terminal sequences.
Install it if you need to display or transform source code.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
See also s3torchconnectorclient · litdata · mosaicml-streaming · webdataset · megatron-fsdp · torchtitan · tensorizer · aws-cdk.aws-kinesisfirehose-alpha · metaflow-checkpoint · s3fs