seqio-nightly
SeqIO: Task-based datasets, preprocessing, and evaluation for sequence models.
Decision gist · record as of 2026-08-14
Yes, if you are building NLP or sequence-model training pipelines and want a structured, reusable way to define tasks and datasets. The library is actively maintained, has low install friction, and is permissively licensed. However, this is a nightly build (0.0.18.dev20250227); verify that the development version is appropriate for your use case. The 11 runtime dependencies add significant install time and disk space.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires TensorFlow and jax as runtime dependencies; initial install of these packages may require compilation or system libraries.
- Low install friction with a pure-Python wheel distribution.
- Active maintenance with recent releases; last commit 2026-07-02.
License · maintenance · safety
Apache 2.0 (permissive) — Apache 2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license notices and may not hold the authors liable.
last release 2025-02-27 (533 days) · last repo commit 2026-07-02 · 595 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 216,891 downloads/mo, #9,370 on PyPI
Alternatives
Verify before relying
pip install seqio-nightly
import seqio
seqio.TaskRegistry.add(
"my_task",
seqio.TfdsDataSource(tfds_name="dataset_name"),
preprocessors=[seqio.preprocessors.tokenize],
output_features={
'inputs': seqio.Feature(seqio.SentencePieceVocabulary('/path/to/vocab')),
'targets': seqio.Feature(seqio.SentencePieceVocabulary('/path/to/vocab'))
}
)
dataset = seqio.get_dataset("my_task")- Whether the nightly build (0.0.18.dev20250227) is suitable for production use versus requiring a stable release
- Specific Python version compatibility, as requires_python is unspecified in the metadata
What it is and what it does
SeqIO is a library for building scalable data pipelines for sequence models, originally extracted from the T5 project. It abstracts the common pattern of combining a data source (TFDS, text files, TFRecord, or custom functions), applying preprocessing steps, defining output features with vocabularies and tokenization, and registering evaluation metrics into named Tasks that can be composed into Mixtures. The library uses tf.data.Dataset internally but is framework-agnostic: you can convert the output to numpy iterators for use with jax, or other frameworks with a single line of code.
SeqIO is designed for sequence-based modalities—text and audio are natural fits, and images work when represented as sequences. It handles the boilerplate of task definition and dataset construction, letting you focus on defining what your data looks like and how to preprocess it. Tasks are typically registered globally so they can be referenced by name in model configs and training scripts.
Use it for
- Build machine translation pipelines combining TFDS datasets with language-pair-specific preprocessing and evaluation metrics
- Define text-to-text tasks with separate input and target features, vocabularies, and custom postprocessors for evaluation
- Create reusable task mixtures that combine multiple datasets for multi-task learning with a single registry lookup
- Preprocess raw text or audio data from files or TFRecord formats into tokenized sequences ready for model training
- Evaluate model outputs using task-specific metrics by defining postprocessors that convert detokenized predictions back to evaluation format
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building NLP or sequence-model training pipelines and want a structured, reusable way to define tasks and datasets.
The library is actively maintained, has low install friction, and is permissively licensed. However, this is a nightly build (0.0.18.dev20250227); verify that the development version is appropriate for your use case. The 11 runtime dependencies add significant install time and disk space.
Install
seqio-nightly on PyPI
Before you install
Low install friction with a pure-Python wheel distribution. Active maintenance with recent releases; last commit 2026-07-02. Depends on 11 runtime packages including jax, tensorflow-text, and tfds-nightly, which may require system-level dependencies or compilation time on first install.
Requires TensorFlow and jax as runtime dependencies; initial install of these packages may require compilation or system libraries.
License in practice
Apache 2.0 permissive license allows commercial and private use with minimal restrictions; you must retain license notices and may not hold the authors liable.
Quickstart
pip install seqio-nightly
import seqio
seqio.TaskRegistry.add(
"my_task",
seqio.TfdsDataSource(tfds_name="dataset_name"),
preprocessors=[seqio.preprocessors.tokenize],
output_features={
'inputs': seqio.Feature(seqio.SentencePieceVocabulary('/path/to/vocab')),
'targets': seqio.Feature(seqio.SentencePieceVocabulary('/path/to/vocab'))
}
)
dataset = seqio.get_dataset("my_task")
Verify before relying
- Whether the nightly build (0.0.18.dev20250227) is suitable for production use versus requiring a stable release
- Specific Python version compatibility, as requires_python is unspecified in the metadata
Package facts
| License | Apache 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 11 packagesabsl-pyclueditdistancejaxjaxlibnumpypackagingpyglovesentencepiecetensorflow-texttfds-nightly |
| Maintenance | Actively maintained 533 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 216,891 / month, #9,370 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseTopic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: seqio_nightly-0.0.18.dev20250227-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “nlp dataset preprocessing”
- seqio-nightlySeqIO provides task-based data pipelines for sequence models,…
- seqioSeqIO builds scalable data pipelines for sequence models using…
- torchtexttorchtext provides text datasets, preprocessing transforms, and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also seqio · tfds-nightly · tensorflow-datasets · tb-nightly · tf2crf · unitxt · torchtext · tensorboard-data-server · keras-hub · tensorflow-text