vega-datasets
A Python package for offline access to Vega datasets
Decision gist · record as of 2026-08-14
Yes, if you need sample datasets for visualization or analysis prototyping and can tolerate that the package is no longer maintained. The low install friction, permissive license, and bundled offline data make it convenient for development and teaching. However, do not rely on it for production pipelines or expect updates to bundled datasets or compatibility with future pandas versions.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with a single pandas dependency and a pure-Python wheel.
- However, the package is archived and abandoned as of 2025-11-12, with no releases since 2020-11-26; it will not receive bug fixes or security updates.
License · maintenance · safety
MIT (permissive) — MIT license is permissive and poses no restrictions on use, modification, or redistribution in commercial or private projects.
last release 2020-11-26 (2087 days) · last repo commit 2025-11-12 · 189 stars · archived
0 known vulnerabilities (OSV.dev, 2026-08-14) · 199,869 downloads/mo, #9,699 on PyPI
Alternatives
Verify before relying
pip install vega_datasets
from vega_datasets import data
df = data.iris()
print(df.head())- Whether bundled datasets remain current with upstream Vega datasets repository
- Compatibility with pandas versions released after 2020-11-26
- Whether HTTP fallback URLs remain valid and accessible
What it is and what it does
vega_datasets is a data-loading utility that gives you quick access to a collection of sample datasets commonly used in data visualization and analysis. It wraps the Vega project's public dataset collection and returns them as Pandas DataFrames. The package bundles a half-dozen datasets (iris, cars, seattle-weather, and others) for offline use, and falls back to fetching additional datasets via HTTP when needed.
You import a single `data` object, call methods like `data.iris()` to get a DataFrame, and optionally inspect the source URL or local file path. The package is designed for exploratory analysis, visualization prototyping, and testing—situations where you need real but small, well-known datasets without setting up a data pipeline.
Use it for
- Prototyping data visualizations with Altair or Matplotlib using well-known reference datasets
- Teaching or demonstrating data analysis workflows with standard, reproducible sample data
- Unit testing visualization or data-processing code with bundled datasets that require no network
- Quick exploratory analysis when you need a familiar dataset like iris or flights without downloading
- Building examples or documentation that reference canonical datasets from the Vega ecosystem
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need sample datasets for visualization or analysis prototyping and can tolerate that the package is no longer maintained.
The low install friction, permissive license, and bundled offline data make it convenient for development and teaching. However, do not rely on it for production pipelines or expect updates to bundled datasets or compatibility with future pandas versions.
Install
vega-datasets on PyPI
Before you install
Low install friction with a single pandas dependency and a pure-Python wheel. However, the package is archived and abandoned as of 2025-11-12, with no releases since 2020-11-26; it will not receive bug fixes or security updates.
License in practice
MIT license is permissive and poses no restrictions on use, modification, or redistribution in commercial or private projects.
Quickstart
pip install vega_datasets
from vega_datasets import data
df = data.iris()
print(df.head())
Verify before relying
- Whether bundled datasets remain current with upstream Vega datasets repository
- Compatibility with pandas versions released after 2020-11-26
- Whether HTTP fallback URLs remain valid and accessible
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.5 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagepandas |
| Maintenance | Abandoned 2,087 days since the last release |
| Last repo commit | repository archived |
| First released | |
| Downloads | 199,869 / month, #9,699 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseNatural Language :: EnglishProgramming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8 |
Evidence: vega_datasets-0.9.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “vega datasets python”
- vega-datasetsProvides offline and online access to Vega datasets as Pandas…
- vegafusionVegaFusion provides Rust-backed tools for analyzing and scaling Vega…
- dvc-renderRenders data in DVC plots format into Vega visualizations and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also palmerpenguins · vegafusion · vl-convert-python · ucimlrepo · datazets · altair · facets-overview · datasets · tfds-nightly · tensorflow-datasets