iden
simple library to manage a dataset of shards to train machine learning models
Decision gist · record as of 2026-08-14
Yes. iden is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a concrete problem in ML workflows—organizing and lazily loading sharded datasets. The permissive BSD-3-Clause license poses no restriction. The API is pre-1.0 and may change, so pin the version if stability is critical, but for new projects or exploratory work it is a solid choice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low friction: pure Python wheel with only two runtime dependencies (coola and objectory).
- Actively maintained with a recent release 60 days ago and current commit activity.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause is permissive; you can use iden in commercial and proprietary projects with minimal restrictions, provided you include the license notice.
last release 2026-06-15 (60 days) · last repo commit 2026-08-14
0 known vulnerabilities (OSV.dev, 2026-08-14) · 583,159 downloads/mo, #5,897 on PyPI
Alternatives
Verify before relying
pip install iden
from iden.shard import create_json_shard
from iden.dataset import create_vanilla_dataset
shard = create_json_shard(data={"key": "value"}, uri="file:///path/to/data.json")
data = shard.get_data()- Performance characteristics when managing large numbers of shards or very large individual shard files.
- Memory overhead of the caching mechanism and how it scales with dataset size.
- Compatibility with distributed training frameworks beyond what the fact sheet documents.
What it is and what it does
iden is a Python library for organizing and accessing machine learning training data split into shards—discrete data chunks that can be stored in different formats and loaded on demand. It abstracts away the mechanics of managing train/validation/test splits, persisting shards to disk, and retrieving them lazily so you don't load everything into memory at once. Each shard has a URI for reproducible identification and optional caching for frequently accessed data.
The library depends on coola (for data comparison) and objectory (for dynamic object instantiation), and supports formats like JSON, YAML, Pickle, PyTorch tensors, and safetensors. It's designed for the common ML workflow where you organize data into logical splits and want to load individual shards on demand rather than materializing the entire dataset upfront.
Use it for
- Organize large training datasets into splits (train/val/test) and load shards lazily during model training.
- Store preprocessed data in multiple formats and switch between them without rewriting shard management code.
- Cache frequently accessed shards in memory while keeping the full dataset on disk to manage memory constraints.
- Persist dataset structure and shard references using URIs for reproducible data pipelines across runs.
- Build custom shard loaders for domain-specific data formats by extending the library's extensible architecture.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
iden is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a concrete problem in ML workflows—organizing and lazily loading sharded datasets. The permissive BSD-3-Clause license poses no restriction. The API is pre-1.0 and may change, so pin the version if stability is critical, but for new projects or exploratory work it is a solid choice.
Install
iden on PyPI
Before you install
Low friction: pure Python wheel with only two runtime dependencies (coola and objectory). Actively maintained with a recent release 60 days ago and current commit activity.
Requires Python 3.10 or later.
License in practice
BSD-3-Clause is permissive; you can use iden in commercial and proprietary projects with minimal restrictions, provided you include the license notice.
Quickstart
pip install iden
from iden.shard import create_json_shard
from iden.dataset import create_vanilla_dataset
shard = create_json_shard(data={"key": "value"}, uri="file:///path/to/data.json")
data = shard.get_data()
Verify before relying
- Performance characteristics when managing large numbers of shards or very large individual shard files.
- Memory overhead of the caching mechanism and how it scales with dataset size.
- Compatibility with distributed training frameworks beyond what the fact sheet documents.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagescoolaobjectory |
| Maintenance | Actively maintained 60 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 583,159 / month, #5,897 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software Development :: Libraries |
Evidence: iden-0.4.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “lazy loading dataset shards”
- ideniden manages machine learning datasets organized into shards with…
- webdatasetWebDataset reads and streams large-scale training data from tar-based…
- datasetdataset simplifies reading and writing data to databases by wrapping…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also webdataset · dataset · litdata · mosaicml-streaming · aiodataloader · datasets · ossdata · coola · ogb · pymongo_inmemory