palmerpenguins
A python package for the palmer penguins dataset
Decision gist · record as of 2026-08-14
Yes, if you need a standard pedagogical dataset for teaching or prototyping. The package is lightweight, dependency-minimal, and carries no security vulnerabilities. Maintenance is aging but the repository remains active. Install it for data exploration tutorials, visualization practice, or as a Iris replacement—not for production data pipelines.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.9
- Low install friction; depends only on pandas and numpy.
- Package is aging (194 days since last release) but the repository remains active and unarchived, suggesting maintenance is minimal but not abandoned.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package freely. The underlying dataset itself is CC-0, placing it in the public domain.
last release 2026-02-01 (194 days) · last repo commit 2026-02-01 · 70 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 133,732 downloads/mo, #11,496 on PyPI
Alternatives
Verify before relying
pip install palmerpenguins
from palmerpenguins import load_penguins
penguins = load_penguins()- Whether the dataset is regularly updated or reflects a fixed snapshot from a specific collection period
- Performance characteristics when working with the full 344-row dataset in typical workflows
What it is and what it does
palmerpenguins is a lightweight data-loading package that wraps the Palmer penguins dataset—a collection of 344 Antarctic penguin observations gathered by Dr. Kristen Gorman and the Palmer Station LTER program. It provides a simple Python interface to load this dataset into a pandas DataFrame, replacing the overused Iris dataset as a standard teaching and exploration tool. The dataset includes size measurements, clutch observations, and blood isotope ratios across three penguin species (Adelie, Chinstrap, Gentoo) observed on islands in the Palmer Archipelago.
The package depends only on pandas and numpy, making it quick to install and integrate into analysis workflows. It is intended for data exploration, visualization practice, and educational purposes rather than production analysis. The underlying data is public domain (CC-0), and the package itself carries a permissive MIT license.
Use it for
- Teaching data visualization and exploratory data analysis to replace the Iris dataset in courses and tutorials
- Prototyping data pipelines and testing pandas workflows with a real, multi-species ecological dataset
- Creating example notebooks and documentation that need a standard, well-documented sample dataset
- Comparing statistical or machine learning techniques across penguin species as a pedagogical exercise
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need a standard pedagogical dataset for teaching or prototyping.
The package is lightweight, dependency-minimal, and carries no security vulnerabilities. Maintenance is aging but the repository remains active. Install it for data exploration tutorials, visualization practice, or as a Iris replacement—not for production data pipelines.
Install
palmerpenguins on PyPI
Before you install
Low install friction; depends only on pandas and numpy. Package is aging (194 days since last release) but the repository remains active and unarchived, suggesting maintenance is minimal but not abandoned.
Requires Python >=3.9
License in practice
MIT license is permissive; you can use, modify, and distribute this package freely. The underlying dataset itself is CC-0, placing it in the public domain.
Quickstart
pip install palmerpenguins
from palmerpenguins import load_penguins
penguins = load_penguins()
Verify before relying
- Whether the dataset is regularly updated or reflects a fixed snapshot from a specific collection period
- Performance characteristics when working with the full 344-row dataset in typical workflows
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespandasnumpy |
| Maintenance | Aging 194 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 133,732 / month, #11,496 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: palmerpenguins-0.1.6-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “palmer penguins dataset”
- palmerpenguinsLoads the Palmer penguins dataset—344 observations of Adelie,…
- mlcroissantValidates and loads datasets described in the MLCommons Croissant…
- athina-clientA Python SDK for logging and interacting with datasets on the Athina…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also vega-datasets · pyLDAvis · datazets · pygwalker · facets-overview · ydata-profiling · ucimlrepo · pandas · feather-format · ibis-framework