kedro-datasets
Kedro-Datasets is where you can find all of Kedro's data connectors.
Decision gist · record as of 2026-08-14
Yes. Kedro-Datasets is actively maintained, has low install friction, carries no known vulnerabilities, and is essential if you're using Kedro for data pipelines. It provides the standard data connectors most projects need and is permissively licensed. Install it as part of your Kedro setup and add optional group dependencies only for the formats your pipeline actually uses.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; optional group-level dependencies (e.g., pandas, spark) must be installed separately for specific dataset types.
- Low friction install with only three runtime dependencies (kedro, backports.strenum, lazy_loader).
- Active maintenance with releases every 7 days and a recent commit history.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache 2.0 (permissive), allowing commercial use and modification with attribution.
last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 118 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,814,647 downloads/mo, #3,524 on PyPI
Alternatives
Verify before relying
pip install kedro-datasets
from kedro_datasets.pandas import CSVDataset
ds = CSVDataset(filepath='data.csv')
df = ds.load()- Whether lazy_loader is used for runtime performance optimization or just module discovery.
- Specific performance characteristics when working with large files or cloud storage.
- Whether all dataset types work equally well across local, network, and cloud backends.
What it is and what it does
Kedro-Datasets is a plugin that extends Kedro's DataCatalog with a collection of AbstractDataset implementations, allowing you to read and write data in many formats (CSV, Excel, Parquet, JSON, HDF5, SQL, Spark, images, and more) across different storage backends (local, network, cloud object stores, Hadoop). It's organized into groups by data type—pandas, spark, networkx, matplotlib, yaml—so you can install only the dependencies you need for your specific workflow.
The package acts as a bridge between Kedro's data abstraction layer and the actual storage and format libraries. Instead of writing custom dataset classes for each file type and storage combination, you declare your data sources in Kedro's configuration and let kedro-datasets handle the loading and saving. It supports optional group-level and dataset-level dependency installation, so a project using only CSV files doesn't need Spark or database drivers.
Use it for
- Load CSV or Parquet files into pandas DataFrames within a Kedro pipeline without writing custom dataset code.
- Read and write Spark DataFrames to cloud object stores (S3, GCS) through Kedro's DataCatalog.
- Query SQL databases and cache results as datasets in a reproducible data pipeline.
- Work with multiple data formats (Excel, JSON, HDF5) in a single Kedro project with consistent APIs.
- Extend Kedro pipelines to handle image data or custom formats by implementing your own AbstractDataset.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Kedro-Datasets is actively maintained, has low install friction, carries no known vulnerabilities, and is essential if you're using Kedro for data pipelines. It provides the standard data connectors most projects need and is permissively licensed. Install it as part of your Kedro setup and add optional group dependencies only for the formats your pipeline actually uses.
Install
kedro-datasets on PyPI
Before you install
Low friction install with only three runtime dependencies (kedro, backports.strenum, lazy_loader). Active maintenance with releases every 7 days and a recent commit history.
Requires Python 3.10 or later; optional group-level dependencies (e.g., pandas, spark) must be installed separately for specific dataset types.
License in practice
Licensed under Apache 2.0 (permissive), allowing commercial use and modification with attribution.
Quickstart
pip install kedro-datasets
from kedro_datasets.pandas import CSVDataset
ds = CSVDataset(filepath='data.csv')
df = ds.load()
Verify before relying
- Whether lazy_loader is used for runtime performance optimization or just module discovery.
- Specific performance characteristics when working with large files or cloud storage.
- Whether all dataset types work equally well across local, network, and cloud backends.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesbackports.strenumkedrolazy_loader |
| Maintenance | Actively maintained 7 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,814,647 / month, #3,524 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: kedro_datasets-9.6.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “kedro data connectors”
- kedro-datasetsKedro-Datasets provides data connectors for Kedro's DataCatalog,…
- kedroKedro is a Python framework for building production-ready data…
- kedro-telemetryKedro-Telemetry is a plugin that collects anonymized usage analytics…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also kedro · kedro-viz · datasets · duckdb-extensions · tablib · pyspark-huggingface · pyspark-extension · dbt-databricks · delta-sharing · parquet-tools