kedro
Kedro helps you build production-ready data and analytics pipelines
Decision gist · record as of 2026-08-14
Yes. Kedro is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache 2.0 license. It is well-suited for teams building reproducible, modular data pipelines at scale. Install it if you need structure and best practices for data engineering or data science projects beyond notebooks.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; the data catalog and deployment integrations may require additional configuration or optional dependencies not bundled with the base package.
- Low install friction with a pure-wheel distribution and 17 well-established runtime dependencies.
- Active maintenance with a recent release 46 days ago and ongoing commits; the project has 10953 stars and is hosted by the LF AI & Data Foundation.
License · maintenance · safety
permissive license (permissive) — Apache 2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions, making it suitable for proprietary and open-source projects alike.
last release 2026-06-29 (46 days) · last repo commit 2026-08-14 · 10,953 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 905,549 downloads/mo, #4,761 on PyPI
Alternatives
Verify before relying
pip install kedro
from kedro.pipeline import Pipeline, node
from kedro.io import DataCatalog
# Define a simple pipeline node
def process_data(input_data):
return input_data * 2
pipeline = Pipeline([node(process_data, "raw_data", "processed_data")])- Whether the data catalog supports all advertised file formats and cloud providers out of the box or if additional plugins are required
- Performance characteristics and scalability limits for large pipelines or distributed execution
- Integration maturity with Argo, Prefect, Kubeflow, AWS Batch, and Databricks mentioned in the description
What it is and what it does
Kedro is a framework for structuring data engineering and data science projects using software engineering best practices. It provides a project template based on Cookiecutter Data Science, a data catalog for managing data connectors across multiple file formats and storage systems, automatic pipeline dependency resolution, and visualization via Kedro-Viz. The framework emphasizes reproducibility, modularity, and team collaboration by moving away from Jupyter notebooks and ad-hoc scripts toward maintainable, reusable code.
The package includes 17 runtime dependencies covering configuration management (OmegaConf, Dynaconf), CLI tooling (Click), templating (Cookiecutter), file system abstraction (fsspec), version control integration (GitPython), and output formatting (Rich). It supports Python 3.10 through 3.14 and is actively maintained by the Kedro product team and open-source contributors.
Use it for
- Build reproducible machine learning pipelines with automatic task dependency tracking and data versioning for file-based systems
- Structure team data science projects with standardized layouts, coding standards, and test-driven development practices
- Deploy data workflows to distributed systems including Argo, Prefect, Kubeflow, AWS Batch, and Databricks
- Manage complex data transformations across multiple file formats and cloud storage backends via a unified data catalog
- Visualize and debug pipeline execution flow and data lineage using Kedro-Viz integration
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Kedro is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache 2.0 license. It is well-suited for teams building reproducible, modular data pipelines at scale. Install it if you need structure and best practices for data engineering or data science projects beyond notebooks.
Install
kedro on PyPI
Before you install
Low install friction with a pure-wheel distribution and 17 well-established runtime dependencies. Active maintenance with a recent release 46 days ago and ongoing commits; the project has 10953 stars and is hosted by the LF AI & Data Foundation.
Requires Python 3.10 or later; the data catalog and deployment integrations may require additional configuration or optional dependencies not bundled with the base package.
License in practice
Apache 2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions, making it suitable for proprietary and open-source projects alike.
Quickstart
pip install kedro
from kedro.pipeline import Pipeline, node
from kedro.io import DataCatalog
# Define a simple pipeline node
def process_data(input_data):
return input_data * 2
pipeline = Pipeline([node(process_data, "raw_data", "processed_data")])
Verify before relying
- Whether the data catalog supports all advertised file formats and cloud providers out of the box or if additional plugins are required
- Performance characteristics and scalability limits for large pipelines or distributed execution
- Integration maturity with Argo, Prefect, Kubeflow, AWS Batch, and Databricks mentioned in the description
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 17 packagesattrsbuildclickcookiecutterdynaconffsspecgitpythonkedro-telemetrymore_itertoolsomegaconfparsepluggyPyYAMLrichtomlitomli-wtyping_extensions |
| Maintenance | Actively maintained 46 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 905,549 / month, #4,761 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: kedro-1.5.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “production data engineering”
- kedroKedro is a Python framework for building production-ready data…
- oil-reservoir-synthesizerGenerates synthetic oil reservoir simulator data using Perlin noise,…
- salabimSalabim is a Python library for discrete event simulation (DES) that…
Give your agent the search over MCP, or paste the wish link into any chat.
More Application Frameworks packages
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Textual is a Python framework for building cross-platform user interfaces that run in the terminal or web browser using a modern, component-based API.
Install it if you're developing CLI tools, dashboards, or interactive terminal applications.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Build and connect to Model Context Protocol servers that expose tools, resources, and prompts to LLM applications over stdio, HTTP, or SSE transports.
Install it if you need to build or connect to servers.
Werkzeug is a WSGI utility library providing request/response objects, URL routing, an interactive debugger, HTTP utilities, and a development server for building web applications.
See also kedro-datasets · kedro-telemetry · kedro-viz · koheesio · kfp · kfp-pipeline-spec · brickflows · kfp-server-api · apache-beam · eventkit