tfds-nightly
tensorflow/datasets is a library of datasets ready to use with TensorFlow.
Decision gist · record as of 2026-08-14
Yes, if you are working with TensorFlow and need standard ML datasets. The library is actively maintained, has low install friction, and carries no known vulnerabilities. However, this is a nightly build (version 4.9.9.dev202510250044)—use the stable release for production unless you specifically need development features. Always verify your right to use each dataset under its own license.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10.
- TensorFlow must be installed separately to use the loaded datasets.
- Low friction installation with a pure-Python wheel.
License · maintenance · safety
Apache 2.0 (permissive) — Apache 2.0 permissive license allows commercial and private use with minimal restrictions. You must retain license notices and may not hold the authors liable. Note the package's own disclaimer: you are responsible for verifying your right to use each dataset under its own license.
last release 2025-10-25 (293 days) · last repo commit 2026-07-29 · 4,581 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 141,749 downloads/mo, #11,233 on PyPI
Alternatives
Verify before relying
pip install tfds-nightly
import tfds_nightly
ds = tfds_nightly.load('mnist', split='train', as_supervised=True)
ds = ds.batch(128).prefetch(10)- Whether this nightly build is suitable for production use or intended only for testing new dataset additions.
- Performance characteristics compared to the stable release, given the development version status.
- Complete list of available datasets in the current nightly version.
What it is and what it does
tfds-nightly is a library that centralizes access to many public machine-learning datasets, downloading and preparing them into standardized tf.data.Dataset objects. It abstracts away the complexity of locating, downloading, and formatting datasets so that standard use cases work immediately—you can load a dataset like MNIST with a single function call and chain it directly into your training pipeline.
The library depends on 18 runtime packages including numpy, pyarrow, protobuf, and tensorflow-metadata to handle data serialization, array operations, and metadata management. It emphasizes simplicity, performance, determinism, and reproducibility: all users get the same examples in the same order, and the library follows best practices for data pipeline efficiency. This is a nightly build, so it tracks development versions of the underlying library.
Use it for
- Load standard ML benchmarks without writing download or preprocessing code.
- Build reproducible training pipelines where dataset order and splits are consistent across runs.
- Prototype models quickly by chaining load output directly into data transformations.
- Access a curated catalog of public datasets with standardized metadata and documentation.
- Integrate dataset loading into workflows with minimal boilerplate.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are working with TensorFlow and need standard ML datasets.
The library is actively maintained, has low install friction, and carries no known vulnerabilities. However, this is a nightly build (version 4.9.9.dev202510250044)—use the stable release for production unless you specifically need development features. Always verify your right to use each dataset under its own license.
Install
tfds-nightly on PyPI
Before you install
Low friction installation with a pure-Python wheel. Active maintenance as of 2026-07-29 with 4581 repository stars. This is a nightly build (version 4.9.9.dev202510250044), so expect development-stage stability; use the stable release for production unless you need cutting-edge features.
Requires Python >=3.10. TensorFlow must be installed separately to use the loaded datasets.
License in practice
Apache 2.0 permissive license allows commercial and private use with minimal restrictions. You must retain license notices and may not hold the authors liable. Note the package's own disclaimer: you are responsible for verifying your right to use each dataset under its own license.
Quickstart
pip install tfds-nightly
import tfds_nightly
ds = tfds_nightly.load('mnist', split='train', as_supervised=True)
ds = ds.batch(128).prefetch(10)
Verify before relying
- Whether this nightly build is suitable for production use or intended only for testing new dataset additions.
- Performance characteristics compared to the stable release, given the development version status.
- Complete list of available datasets in the current nightly version.
Package facts
| License | Apache 2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 18 packagesabsl-pyarray_recorddm-treeetilsimmutabledictnumpypromiseprotobufpsutilpyarrowrequestssimple_parsingtensorflow-metadatatermcolortomltqdmwraptimportlib_resources |
| Maintenance | Actively maintained 293 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 141,749 / month, #11,233 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyTopic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: tfds_nightly-4.9.9.dev202510250044-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ml dataset download and prepare”
- tfds-nightlyProvides a library of ready-to-use public datasets formatted as…
- ogbOGB provides standardized benchmark datasets, data loaders, and…
- tensorflow-datasetsProvides access to many public datasets as tf.data.Datasets, handling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also tensorflow-datasets · sodapy · seqio-nightly · seqio · opendatalab · tensorflow-metadata · datasets · tb-nightly · tensorboard-data-server · ir-datasets