pyjanitor
Tools for cleaning pandas DataFrames
Decision gist · record as of 2026-08-14
Yes. pyjanitor solves a real readability and usability gap in pandas workflows. Low install friction, active maintenance, permissive MIT license, no known vulnerabilities, and strong adoption make it a safe, practical choice for teams that value readable data-cleaning code. Install it if you work with pandas and want method chains to replace imperative preprocessing logic.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8 or later; janitor-rs (a Rust-based dependency) must compile on your platform.
- Low friction install with five runtime dependencies.
- Actively maintained—last commit 2026-08-11, 1500 repository stars.
License · maintenance · safety
MIT (permissive) — MIT license (permissive) places no restrictions on use, modification, or distribution in proprietary or open-source projects.
last release 2026-04-07 (129 days) · last repo commit 2026-08-11 · 1,500 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 383,172 downloads/mo, #7,079 on PyPI
Alternatives
Verify before relying
import pandas as pd
import janitor
df = pd.DataFrame({'Name': ['Alice', 'Bob'], 'Age': [None]})
df = df.clean_names().remove_empty()- Performance characteristics when chaining many operations on large DataFrames.
- Compatibility of janitor-rs compilation on all target platforms (Windows, macOS, Linux variants).
- Extent of experimental submodules (finance, biology, chemistry, engineering, pyspark) and their stability.
What it is and what it does
pyjanitor is a pandas extension that adds a collection of data-cleaning methods designed to work as part of method chains. Instead of writing imperative pandas code with intermediate variable assignments, you can express a sequence of cleaning steps as a readable chain of method calls—each with an explicit verb name like clean_names(), remove_empty(), or rename_column(). The package is inspired by the R janitor package and the dplyr paradigm, bringing that fluent, declarative style to Python.
Under the hood, pyjanitor registers its functions as pandas DataFrame methods via pandas-flavor and depends on scipy, natsort, multipledispatch, and janitor-rs for its operations. It handles common preprocessing tasks: standardizing column names, removing null or empty rows and columns, identifying duplicates, encoding categorical data, splitting features and targets, coalescing columns, date conversions, and expanding delimited values into dummy variables. The package is actively maintained, supports Python 3.8 and later, and carries no known security vulnerabilities.
Use it for
- Clean messy column names (spaces, mixed case, special characters) in a single method call before analysis.
- Chain multiple DataFrame transformations (drop nulls, rename columns, add computed fields) in a single readable expression.
- Preprocess raw data exports (Excel, CSV) by removing empty rows and columns and standardizing formats in one pipeline.
- Prepare machine learning datasets by splitting features and targets and encoding categorical variables declaratively.
- Expand delimited or categorical columns into dummy-encoded variables for statistical modeling.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
pyjanitor solves a real readability and usability gap in pandas workflows. Low install friction, active maintenance, permissive MIT license, no known vulnerabilities, and strong adoption make it a safe, practical choice for teams that value readable data-cleaning code. Install it if you work with pandas and want method chains to replace imperative preprocessing logic.
Install
pyjanitor on PyPI
Before you install
Low friction install with five runtime dependencies. Actively maintained—last commit 2026-08-11, 1500 repository stars. Supports Python 3.8 and later.
Requires Python 3.8 or later; janitor-rs (a Rust-based dependency) must compile on your platform.
License in practice
MIT license (permissive) places no restrictions on use, modification, or distribution in proprietary or open-source projects.
Quickstart
import pandas as pd
import janitor
df = pd.DataFrame({'Name': ['Alice', 'Bob'], 'Age': [None]})
df = df.clean_names().remove_empty()
Verify before relying
- Performance characteristics when chaining many operations on large DataFrames.
- Compatibility of janitor-rs compilation on all target platforms (Windows, macOS, Linux variants).
- Extent of experimental submodules (finance, biology, chemistry, engineering, pyspark) and their stability.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesnatsortpandas-flavormultipledispatchscipyjanitor-rs |
| Maintenance | Actively maintained 129 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 383,172 / month, #7,079 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/Engineering |
Evidence: pyjanitor-0.32.23-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pandas data cleaning methods”
- pyjanitorExtends pandas DataFrames with method-chainable data-cleaning…
- pygwalkerTurns pandas DataFrames into interactive drag-and-drop visual…
- skrubskrub prepares and transforms dataframes for machine learning by…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also pandas-flavor · skrub · gspread-pandas · gspread-dataframe · awkward-pandas · pygwalker · swifter · banal · pandas-read-xml