$npx skillfedfor your agent

seqio

SeqIO: Task-based datasets, preprocessing, and evaluation for sequence models.

With conditionsPyPI Artificial IntelligenceReleased Aug 2025301.4K downloads / moApache 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — seqio-0.0.20-py2.py3-none-any.whl
v0.0.20 · released 2025-08-28 · 11 runtime deps: absl-py, clu, editdistance, jax, jaxlib, numpy, packaging, pyglove

Yes, if you are building sequence models (NLP, audio, or sequence-based vision) and want a structured, reusable way to manage data pipelines with built-in preprocessing and evaluation. The low install friction and active maintenance support this. However, the 11 runtime dependencies—particularly TensorFlow, JAX, and TensorFlow Text—make it heavy for lightweight use cases; consider it only if you need task-based dataset abstraction and don't already have a simpler pipeline in place.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires TensorFlow, JAX, and TensorFlow Text as runtime dependencies; compatible with other frameworks via numpy iterator conversion but TensorFlow installation is mandatory.
  • Low friction installation with a pure-Python wheel.
  • The package is actively maintained with recent releases and moderate popularity (top_15000 tier), though it carries 11 runtime dependencies including JAX, TensorFlow, and TensorFlow Text, which may add setup complexity in constrained environments.

License · maintenance · safety

Apache 2.0 (permissive) — Licensed under Apache 2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

last release 2025-08-28 (351 days) · last repo commit 2026-07-02 · 595 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 301,423 downloads/mo, #7,837 on PyPI

Verify before relying

pip install seqio

import seqio
import tensorflow as tf

task = seqio.TaskRegistry.add(
    "example_task",
    seqio.TfdsDataSource(tfds_name="dataset_name"),
    output_features={
        'inputs': seqio.Feature(seqio.PassThroughVocabulary(), dtype=tf.int32),
        'targets': seqio.Feature(seqio.PassThroughVocabulary(), dtype=tf.int32),
    }
)

dataset = seqio.get_dataset("example_task", batch_size=32)
  • Minimum Python version requirement (requires_python is unspecified in metadata)
  • Whether all 11 runtime dependencies are truly required for basic usage or if some are optional
  • Performance characteristics and scalability limits for very large datasets
Same gist for agents: .md · .json

What it is and what it does

SeqIO is a library for constructing data pipelines for sequence models, built on top of TensorFlow's tf.data.Dataset. It abstracts away the complexity of loading raw data, applying preprocessing steps, tokenizing features with custom vocabularies, and computing evaluation metrics into a unified Task interface. The library was originally extracted from the T5 model's data pipeline and refactored for general use.

The package is designed to work primarily with sequential data—text and audio are naturally supported, and images can be used if represented as sequences. While it uses TensorFlow internally, SeqIO can output datasets as numpy iterators, making it fully compatible with JAX, PyTorch, and other frameworks. You define a Task by specifying a data source (TFDS, text files, TFRecord, or custom functions), preprocessing steps, output feature definitions with vocabularies, and metric functions, then use seqio.get_dataset to obtain a ready-to-use tf.data.Dataset.

Use it for

  • Building machine translation pipelines with preprocessing and BLEU evaluation for sequence-to-sequence models
  • Creating text-to-text task datasets with custom tokenization and prompt formatting for transfer learning
  • Preprocessing benchmark datasets from TensorFlow Datasets with task-specific metrics for model evaluation
  • Combining multiple tasks into a Mixture for multi-task learning with unified data handling
  • Converting raw text or audio files into tokenized sequences with vocabulary management for downstream models

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are building sequence models (NLP, audio, or sequence-based vision) and want a structured, reusable way to manage data pipelines with built-in preprocessing and evaluation.

The low install friction and active maintenance support this. However, the 11 runtime dependencies—particularly TensorFlow, JAX, and TensorFlow Text—make it heavy for lightweight use cases; consider it only if you need task-based dataset abstraction and don't already have a simpler pipeline in place.

Install

seqio on PyPI

Before you install

Low friction installation with a pure-Python wheel. The package is actively maintained with recent releases and moderate popularity (top_15000 tier), though it carries 11 runtime dependencies including JAX, TensorFlow, and TensorFlow Text, which may add setup complexity in constrained environments.

Requires TensorFlow, JAX, and TensorFlow Text as runtime dependencies; compatible with other frameworks via numpy iterator conversion but TensorFlow installation is mandatory.

License in practice

Licensed under Apache 2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

Quickstart

pip install seqio

import seqio
import tensorflow as tf

task = seqio.TaskRegistry.add(
    "example_task",
    seqio.TfdsDataSource(tfds_name="dataset_name"),
    output_features={
        'inputs': seqio.Feature(seqio.PassThroughVocabulary(), dtype=tf.int32),
        'targets': seqio.Feature(seqio.PassThroughVocabulary(), dtype=tf.int32),
    }
)

dataset = seqio.get_dataset("example_task", batch_size=32)

Verify before relying

  • Minimum Python version requirement (requires_python is unspecified in metadata)
  • Whether all 11 runtime dependencies are truly required for basic usage or if some are optional
  • Performance characteristics and scalability limits for very large datasets

Package facts

LicenseApache 2.0 permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
11 packages
absl-pyclueditdistancejaxjaxlibnumpypackagingpyglovesentencepiecetensorflow-texttensorflow-datasets
MaintenanceActively maintained 351 days since the last release
Last repo commit
First released
Downloads301,423 / month, #7,837 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseTopic :: Scientific/Engineering :: Artificial Intelligence

Evidence: seqio-0.0.20-py2.py3-none-any.whl

Tags

Capabilities
sequence data preprocessing pipelinenlp dataset preprocessingmachine learning data pipelinetext tokenization and evaluationtask-based dataset managementtensorflow data pipelinesequence model training data
Topics
nlp-data-pipelinetensorflow-ecosystemsequence-models
PyPI keywords
sequencepreprocessingnlpmachinelearning

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “sequence data preprocessing pipeline”

  • seqioSeqIO builds scalable data pipelines for sequence models using…
  • seqio-nightlySeqIO provides task-based data pipelines for sequence models,…
  • Keras-PreprocessingProvides data preprocessing and augmentation utilities for deep…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also seqio-nightly · tensorflow-datasets · tfds-nightly · tensorflow-text · Keras-Preprocessing · datasets · torchtext · unitxt · keras-nlp · tensorflow-recommenders