woodwork
a data typing library for machine learning
Decision gist · record as of 2026-08-14
Yes, if you are using Featuretools or EvalML and need a standardized way to annotate and manage DataFrame column types. The low install friction and permissive license make it a straightforward addition. However, be aware that the project is aging—last release was May 2024 and there have been no commits since September 2025—so verify compatibility with your current versions of pandas and scikit-learn before relying on it for new projects.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with 6 common runtime dependencies (pandas, scikit-learn, scipy, numpy, python-dateutil, importlib-resources).
- Maintenance status is aging—last release was in May 2024 and no commits since September 2025, suggesting the project is stable but not actively developed.
License · maintenance · safety
permissive license (permissive) — BSD 3-Clause License is permissive; you can use, modify, and distribute Woodwork freely in commercial and private projects as long as you retain the license notice and disclaimer.
last release 2024-05-14 (822 days) · last repo commit 2025-09-30 · 155 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 170,445 downloads/mo, #10,394 on PyPI
Alternatives
Verify before relying
import pandas as pd
import woodwork as ww
df = pd.read_csv('data.csv')
df.ww.init(name='my_data')
df.ww.set_types(logical_types={'col_name': 'Integer', 'category_col': 'Categorical'})
filtered = df.ww.select(include=['numeric'])- Whether the aging maintenance status (last release May 2024, no commits since Sept 2025) affects compatibility with recent pandas or scikit-learn versions.
- Whether Woodwork's integration with Featuretools and EvalML remains current given the maintenance gap.
What it is and what it does
Woodwork is a data typing library that extends pandas DataFrames with a semantic layer for machine learning. It lets you assign logical types (Integer, Categorical, DateTime, NaturalLanguage, etc.) and semantic tags to columns, then query and filter your data based on those annotations. The library automatically infers types from underlying data when you don't specify them, and it stores metadata alongside your DataFrame for use in downstream ML workflows.
The package is designed as a common typing namespace for Featuretools and EvalML, so if you're using those tools, Woodwork provides a standardized way to communicate data structure and meaning. It depends on pandas, scikit-learn, scipy, numpy, python-dateutil, and importlib-resources—all standard data-science libraries—so installation is straightforward. The project is maintained by Alteryx but has not seen active development since mid-2024.
Use it for
- Annotate DataFrame columns with logical types and semantic tags before passing data to Featuretools for automated feature engineering.
- Store and communicate column metadata (e.g., which columns are numeric, categorical, or contain personal names) for reproducible ML pipelines.
- Filter and select DataFrame subsets by logical type or semantic tag without manual column name lists.
- Standardize data typing across multiple DataFrames in a machine learning project to ensure consistent downstream processing.
- Integrate with EvalML for automated machine learning workflows that rely on consistent column type information.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are using Featuretools or EvalML and need a standardized way to annotate and manage DataFrame column types.
The low install friction and permissive license make it a straightforward addition. However, be aware that the project is aging—last release was May 2024 and there have been no commits since September 2025—so verify compatibility with your current versions of pandas and scikit-learn before relying on it for new projects.
Install
woodwork on PyPI
Before you install
Low install friction with 6 common runtime dependencies (pandas, scikit-learn, scipy, numpy, python-dateutil, importlib-resources). Maintenance status is aging—last release was in May 2024 and no commits since September 2025, suggesting the project is stable but not actively developed.
License in practice
BSD 3-Clause License is permissive; you can use, modify, and distribute Woodwork freely in commercial and private projects as long as you retain the license notice and disclaimer.
Quickstart
import pandas as pd
import woodwork as ww
df = pd.read_csv('data.csv')
df.ww.init(name='my_data')
df.ww.set_types(logical_types={'col_name': 'Integer', 'category_col': 'Categorical'})
filtered = df.ww.select(include=['numeric'])
Verify before relying
- Whether the aging maintenance status (last release May 2024, no commits since Sept 2025) affects compatibility with recent pandas or scikit-learn versions.
- Whether Woodwork's integration with Featuretools and EvalML remains current given the maintenance gap.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release <4,>=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagespandasscikit-learnpython-dateutilscipyimportlib-resourcesnumpy |
| Maintenance | Aging 822 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 170,445 / month, #10,394 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: MacOSOperating System :: Microsoft :: WindowsOperating System :: POSIXOperating System :: UnixProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Software Development |
Evidence: woodwork-0.31.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “dataframe column typing”
- woodworkWoodwork adds a typing layer to pandas DataFrames, letting you…
- quinnQuinn provides helper methods for PySpark DataFrame validation,…
- semantic-link-functions-phonenumbersAdds phone number validation as a semantic function to…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also semantic-link-functions-holidays · visions · semantic-link-functions-phonenumbers · sklearn-pandas · awkward-pandas · pandas-summary · featuretools · Pint-Pandas · pandas · datacompy