tensorflow-transform
A library for data preprocessing with TensorFlow
Decision gist · record as of 2026-08-14
Yes. TensorFlow Transform is production-stable, actively maintained, and essential for building ML pipelines that require full-pass preprocessing or consistent train-serve transformations. Install if you need stateful data transformations (normalization, vocabulary generation, bucketing) in TensorFlow workflows; skip if your preprocessing fits single-example operations.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires TensorFlow, Apache Beam, and PyArrow as runtime dependencies; Python 3.10 or later.
- Low friction installation as a pure-Python wheel.
- Active maintenance with a recent release (64 days old) and ongoing repository activity.
License · maintenance · safety
Apache 2.0 (permissive) — Licensed under Apache 2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
last release 2026-06-11 (64 days) · last repo commit 2026-08-14 · 988 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 187,343 downloads/mo, #9,968 on PyPI
Alternatives
Verify before relying
pip install tensorflow-transform
import tensorflow_transform as tft
# Use tft.scale_to_z_score() or tft.compute_and_apply_vocabulary()
# within a tf.Transform preprocessing function- Whether the package works with Google Cloud Dataflow or other Apache Beam runners without additional configuration.
- Performance characteristics and scalability limits for very large datasets.
- Compatibility with custom Apache Beam runners beyond the default local mode.
What it is and what it does
TensorFlow Transform is a preprocessing library that extends TensorFlow's single-example capabilities to support full-pass operations over entire datasets. It handles transformations that require seeing all data—such as computing mean and standard deviation for normalization, building vocabularies from all unique values, or assigning data to quantile-based buckets—then exports the result as a frozen TensorFlow graph for reuse in both training and serving pipelines.
The library uses Apache Beam for distributed computation (defaulting to local mode but supporting Google Cloud Dataflow and other runners) and PyArrow for efficient vectorized operations. By applying identical transformations at train and serve time, it eliminates training-serving skew, a common source of model degradation in production.
Use it for
- Normalize numerical features by computing global mean and standard deviation across the entire training dataset.
- Generate a vocabulary from all unique string values in a column and map strings to integer IDs consistently.
- Assign continuous values to discrete buckets based on observed data quantiles or distribution.
- Preprocess data pipelines for TensorFlow Extended (TFX) workflows with distributed Apache Beam runners.
- Export preprocessing logic as a reusable TensorFlow graph to serve alongside trained models.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
TensorFlow Transform is production-stable, actively maintained, and essential for building ML pipelines that require full-pass preprocessing or consistent train-serve transformations. Install if you need stateful data transformations (normalization, vocabulary generation, bucketing) in TensorFlow workflows; skip if your preprocessing fits single-example operations.
Install
tensorflow-transform on PyPI
Before you install
Low friction installation as a pure-Python wheel. Active maintenance with a recent release (64 days old) and ongoing repository activity. Requires TensorFlow and Apache Beam as core dependencies.
Requires TensorFlow, Apache Beam, and PyArrow as runtime dependencies; Python 3.10 or later.
License in practice
Licensed under Apache 2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
Quickstart
pip install tensorflow-transform
import tensorflow_transform as tft
# Use tft.scale_to_z_score() or tft.compute_and_apply_vocabulary()
# within a tf.Transform preprocessing function
Verify before relying
- Whether the package works with Google Cloud Dataflow or other Apache Beam runners without additional configuration.
- Performance characteristics and scalability limits for very large datasets.
- Compatibility with custom Apache Beam runners beyond the default local mode.
Package facts
| License | Apache 2.0 permissive |
| Python support | Supports the current Python release <4,>=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 10 packagesabsl-pyapache-beamnumpyprotobufpyarrowpydottensorflowtensorflow-metadatatf_kerastfx-bsl |
| Maintenance | Actively maintained 64 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 187,343 / month, #9,968 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: MathematicsTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: tensorflow_transform-1.21.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “tensorflow data preprocessing”
- tensorflow-transformTensorFlow Transform preprocesses data with full-pass operations like…
- seqioSeqIO builds scalable data pipelines for sequence models using…
- tensorflow-textTensorFlow Text provides text preprocessing operations and tokenizers…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also tensorflow-data-validation · tfx-bsl · tensorflow-metadata · tensorflow-io · fasttransform · tensorflow-io-gcs-filesystem · diastatic-malt · Keras-Preprocessing · tensorflow-text · tensorflow