$npx skillfedfor your agent

percentify

Data Exploratory stats and Quality diagnostics for pandas and Polars DataFrames. One easy call each.

Worth itPyPI Information AnalysisReleased Jul 2026191.0K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — percentify-1.0.2-py3-none-any.whl
v1.0.2 · released 2026-07-13 · Python >=3.10 · 3 runtime deps: numpy, pandas, scipy

Yes. Percentify fills a real gap: it automates the 80% of exploratory checks you run on every dataset into one or two function calls, with low install friction, active maintenance, permissive licensing, and genuine dual support for pandas and Polars. No known vulnerabilities. Best for teams doing frequent data intake and quality validation.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later and pandas 2.0+.
  • Low install friction with a pure-Python wheel and three well-established runtime dependencies (numpy, pandas, scipy).
  • Active maintenance with a release 32 days ago.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 is permissive; you can use this in commercial and proprietary projects with minimal restrictions, provided you include a copy of the license.

last release 2026-07-13 (32 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 190,981 downloads/mo, #9,896 on PyPI

Verify before relying

pip install percentify

import pandas as pd
from percentify import profiler

df = pd.DataFrame({"salary": [50000, None, 60000], "age": [25, 30, None]})
report = profiler(df)
print(report.to_frame())  # ranked findings with fixes
print(report.health)      # 0-100 data-health score
  • Whether the package's Polars support is truly feature-parity with pandas across all functions.
  • Performance characteristics on large datasets (memory usage, runtime scaling).
  • Availability and completeness of the documentation at data-centt.github.io/percentify.
Same gist for agents: .md · .json

What it is and what it does

Percentify is a data diagnostics library that wraps common exploratory and quality-check operations into single-call functions for pandas and Polars DataFrames. Its flagship function, `profiler()`, scans a DataFrame for data issues—missing values, outliers, collinearity, class imbalance, skew—ranks them by severity, and suggests fixes. The library also provides functions for variance analysis (cv, pca_variance, pca_loadings), statistical testing (permutation_test, bootstrap_ci, effect_size), correlation detection, and formatting utilities.

The package is designed around a specific philosophy: each function returns the single most common answer in one call, with clear output sorted worst-first, and points users to underlying libraries (pandas, scipy, statsmodels, scikit-learn) when deeper customization is needed. It treats pandas and Polars as first-class backends—pass either type and get the same type back without flags or manual conversion. Runtime dependencies are numpy, pandas, and scipy.

Use it for

  • Run `profiler()` before modeling to catch data issues and get a 0-100 health score for a CI data-quality gate.
  • Use `missing()` to quickly see which columns have gaps and how much data is lost per column.
  • Call `vif()` to detect multicollinearity and identify which features to drop before regression.
  • Apply `outliers()` to find the percentage of outliers in each column and decide on treatment.
  • Use `correlate()` to find feature pairs that move together and test whether the relationship is statistically significant.
  • Call `imbalance()` on a target column to measure class skew and decide on resampling or weighting strategies.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Percentify fills a real gap: it automates the 80% of exploratory checks you run on every dataset into one or two function calls, with low install friction, active maintenance, permissive licensing, and genuine dual support for pandas and Polars. No known vulnerabilities. Best for teams doing frequent data intake and quality validation.

Install

percentify on PyPI

Before you install

Low install friction with a pure-Python wheel and three well-established runtime dependencies (numpy, pandas, scipy). Active maintenance with a release 32 days ago.

Requires Python 3.10 or later and pandas 2.0+.

License in practice

Apache-2.0 is permissive; you can use this in commercial and proprietary projects with minimal restrictions, provided you include a copy of the license.

Quickstart

pip install percentify

import pandas as pd
from percentify import profiler

df = pd.DataFrame({"salary": [50000, None, 60000], "age": [25, 30, None]})
report = profiler(df)
print(report.to_frame())  # ranked findings with fixes
print(report.health)      # 0-100 data-health score

Verify before relying

  • Whether the package's Polars support is truly feature-parity with pandas across all functions.
  • Performance characteristics on large datasets (memory usage, runtime scaling).
  • Availability and completeness of the documentation at data-centt.github.io/percentify.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
numpypandasscipy
MaintenanceActively maintained 32 days since the last release
First released
Downloads190,981 / month, #9,896 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13

Evidence: percentify-1.0.2-py3-none-any.whl

Tags

Capabilities
data quality profiling pandas polarsexploratory data analysis one functiondetect outliers missing values collinearitydata health score diagnosticseda statistics automated checkspandas dataframe profilermulticollinearity vif detection
Topics
data-qualityexploratory-analysispandas-polars
PyPI keywords
data-sciencestatisticspandasedamulticollinearityvifpcaoutliersdata-analysis

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “data quality profiling pandas polars”

  • percentifyPercentify provides one-call exploratory statistics and data-quality…
  • cualleeCuallee provides a dataframe-agnostic API to define and run data…
  • datacompyDataComPy compares two DataFrames across Pandas, Polars, Spark, and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also pandas-profiling · datacompy · narwhals · pandas-summary · itables · prince · dataframe-api-compat · ydata-profiling · mrmr-selection · great-tables