$npx skillfedfor your agent

kedro-datasets

Kedro-Datasets is where you can find all of Kedro's data connectors.

Worth itPyPI DatabaseReleased Aug 20261.8M downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — kedro_datasets-9.6.0-py3-none-any.whl
v9.6.0 · released 2026-08-07 · Python >=3.10 · 3 runtime deps: backports.strenum, kedro, lazy_loader

Yes. Kedro-Datasets is actively maintained, has low install friction, carries no known vulnerabilities, and is essential if you're using Kedro for data pipelines. It provides the standard data connectors most projects need and is permissively licensed. Install it as part of your Kedro setup and add optional group dependencies only for the formats your pipeline actually uses.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; optional group-level dependencies (e.g., pandas, spark) must be installed separately for specific dataset types.
  • Low friction install with only three runtime dependencies (kedro, backports.strenum, lazy_loader).
  • Active maintenance with releases every 7 days and a recent commit history.

License · maintenance · safety

permissive license (permissive) — Licensed under Apache 2.0 (permissive), allowing commercial use and modification with attribution.

last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 118 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,814,647 downloads/mo, #3,524 on PyPI

Verify before relying

pip install kedro-datasets

from kedro_datasets.pandas import CSVDataset

ds = CSVDataset(filepath='data.csv')
df = ds.load()
  • Whether lazy_loader is used for runtime performance optimization or just module discovery.
  • Specific performance characteristics when working with large files or cloud storage.
  • Whether all dataset types work equally well across local, network, and cloud backends.
Same gist for agents: .md · .json

What it is and what it does

Kedro-Datasets is a plugin that extends Kedro's DataCatalog with a collection of AbstractDataset implementations, allowing you to read and write data in many formats (CSV, Excel, Parquet, JSON, HDF5, SQL, Spark, images, and more) across different storage backends (local, network, cloud object stores, Hadoop). It's organized into groups by data type—pandas, spark, networkx, matplotlib, yaml—so you can install only the dependencies you need for your specific workflow.

The package acts as a bridge between Kedro's data abstraction layer and the actual storage and format libraries. Instead of writing custom dataset classes for each file type and storage combination, you declare your data sources in Kedro's configuration and let kedro-datasets handle the loading and saving. It supports optional group-level and dataset-level dependency installation, so a project using only CSV files doesn't need Spark or database drivers.

Use it for

  • Load CSV or Parquet files into pandas DataFrames within a Kedro pipeline without writing custom dataset code.
  • Read and write Spark DataFrames to cloud object stores (S3, GCS) through Kedro's DataCatalog.
  • Query SQL databases and cache results as datasets in a reproducible data pipeline.
  • Work with multiple data formats (Excel, JSON, HDF5) in a single Kedro project with consistent APIs.
  • Extend Kedro pipelines to handle image data or custom formats by implementing your own AbstractDataset.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Kedro-Datasets is actively maintained, has low install friction, carries no known vulnerabilities, and is essential if you're using Kedro for data pipelines. It provides the standard data connectors most projects need and is permissively licensed. Install it as part of your Kedro setup and add optional group dependencies only for the formats your pipeline actually uses.

Install

kedro-datasets on PyPI

Before you install

Low friction install with only three runtime dependencies (kedro, backports.strenum, lazy_loader). Active maintenance with releases every 7 days and a recent commit history.

Requires Python 3.10 or later; optional group-level dependencies (e.g., pandas, spark) must be installed separately for specific dataset types.

License in practice

Licensed under Apache 2.0 (permissive), allowing commercial use and modification with attribution.

Quickstart

pip install kedro-datasets

from kedro_datasets.pandas import CSVDataset

ds = CSVDataset(filepath='data.csv')
df = ds.load()

Verify before relying

  • Whether lazy_loader is used for runtime performance optimization or just module discovery.
  • Specific performance characteristics when working with large files or cloud storage.
  • Whether all dataset types work equally well across local, network, and cloud backends.

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
backports.strenumkedrolazy_loader
MaintenanceActively maintained 7 days since the last release
Last repo commit
First released
Downloads1,814,647 / month, #3,524 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: kedro_datasets-9.6.0-py3-none-any.whl

Tags

Capabilities
kedro data connectorsdata catalog abstractioncsv parquet excel datasetsspark sql data loadingcloud storage data accesskedro plugin datasetsfile format data adapters
Topics
data-pipelinekedro-pluginmulti-format

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “kedro data connectors”

  • kedro-datasetsKedro-Datasets provides data connectors for Kedro's DataCatalog,…
  • kedroKedro is a Python framework for building production-ready data…
  • kedro-telemetryKedro-Telemetry is a plugin that collects anonymized usage analytics…

Give your agent the search over MCP, or paste the wish link into any chat.

More Database packages

psycopg2-binary Worth it
PyPI · Software Development · released Apr 2026

psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.

copyleftcompiled wheel · 3.9+
271.6Mdownloads / mo
redis Worth it
PyPI · Database · released Jul 2026

Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.

Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.

MITpure Python · 3.10+
268.3Mdownloads / mo
ydb Worth it
PyPI · Database · released Jul 2026

YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.

Install it if you need to connect Python applications to YDB databases.

permissive licensepure Python · 3.10+
210.0Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
sqlparse Worth it
PyPI · Software Development · released Aug 2026

sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.

Install it if you need to manipulate, format, or analyze SQL text programmatically.

BSD-3-Clausepure Python · 3.10+
148.9Mdownloads / mo
dbt-adapters With conditions
PyPI · Database · released Jul 2026

Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.

Apache-2.0pure Python · 3.10.0+
121.3Mdownloads / mo

See also kedro · kedro-viz · datasets · duckdb-extensions · tablib · pyspark-huggingface · pyspark-extension · dbt-databricks · delta-sharing · parquet-tools