dataengine
General purpose data engineering python package.
Decision gist · record as of 2026-08-14
Yes, with conditions. Install if you need a lightweight configuration-driven wrapper around Spark SQL pipelines and have the Spark/Java runtime available. The low install friction and lack of known vulnerabilities are positives. However, the Alpha status, 484-day staleness, unclear license, and heavy dependency footprint (18 runtime deps) mean you should verify the license terms, confirm the configuration format works for your use case (the docs have a TODO), and assess whether the project's maintenance cadence suits your risk tolerance.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.9.
- Apache Spark and its Java dependencies must be available in your environment for full functionality.
- Low install friction with a pure-Python wheel, but the package is in Alpha status and has not been updated in 484 days.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text is available in the package metadata. Before adopting this in a commercial or regulated context, verify the actual license terms directly from the repository.
last release 2025-04-17 (484 days) · last repo commit 2025-04-17
0 known vulnerabilities (OSV.dev, 2026-08-14) · 159,859 downloads/mo, #10,687 on PyPI
Alternatives
Verify before relying
pip install dataengine
from dataengine import Engine
# Configure Engine with Database, Dataset, and Query subclasses
engine = Engine()
# (See repository for configuration file format)- What configuration file format does Engine expect, and are there working examples beyond the TODO note in the description?
- Does the package support Python versions beyond 3.9, or is 3.9 the only tested version?
- What is the actual license of this package, and under what terms can it be used?
What it is and what it does
Dataengine is a configuration-driven data pipeline framework that wraps Apache Spark, pandas, and cloud service SDKs (AWS, Databricks, GitHub, Slack, Datadog) into a unified Python interface. The core abstraction is the Engine class, which orchestrates three main components: Database objects that represent data stores you connect to, Dataset objects that define data sources (local or S3), and Query objects that specify SQL transformations, input dependencies, and output destinations.
The package is designed to let you define complex data workflows declaratively through configuration files rather than imperative code. It sits in Alpha status and has not been updated in 484 days, suggesting either stable maintenance or dormancy. With 18 runtime dependencies including pyspark, pandas, boto3, and various cloud SDKs, it brings a large dependency footprint and assumes you have Spark and Java available.
Use it for
- Define multi-step SQL transformations on Spark DataFrames with input/output dependencies managed through configuration.
- Orchestrate data loading from S3 or local sources, apply transformations, and write results back to databases or cloud storage.
- Integrate Slack notifications, GitHub metadata, or Datadog metrics into data pipeline workflows via SDK support.
- Manage multiple database connections (PostgreSQL, MySQL, Databricks) from a single Engine instance configured declaratively.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Install if you need a lightweight configuration-driven wrapper around Spark SQL pipelines and have the Spark/Java runtime available. The low install friction and lack of known vulnerabilities are positives. However, the Alpha status, 484-day staleness, unclear license, and heavy dependency footprint (18 runtime deps) mean you should verify the license terms, confirm the configuration format works for your use case (the docs have a TODO), and assess whether the project's maintenance cadence suits your risk tolerance.
Install
dataengine on PyPI
Before you install
Low install friction with a pure-Python wheel, but the package is in Alpha status and has not been updated in 484 days. It carries 18 runtime dependencies including heavy libraries like pyspark and pandas, which will pull in substantial transitive requirements.
Requires Python >=3.9. Apache Spark and its Java dependencies must be available in your environment for full functionality.
License in practice
License status is unclear—no SPDX identifier or raw license text is available in the package metadata. Before adopting this in a commercial or regulated context, verify the actual license terms directly from the repository.
Quickstart
pip install dataengine
from dataengine import Engine
# Configure Engine with Database, Dataset, and Query subclasses
engine = Engine()
# (See repository for configuration file format)
Verify before relying
- What configuration file format does Engine expect, and are there working examples beyond the TODO note in the description?
- Does the package support Python versions beyond 3.9, or is 3.9 the only tested version?
- What is the actual license of this package, and under what terms can it be used?
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 18 packagespytestpytest-covPyYAMLnumpypandaspyarrowmotoboto3psycopg2-binaryPyMySQLslack-sdktabulatedatabricks-cliPyGithubscipymarshmallowdatadog_api_clientpyspark |
| Maintenance | Aging 484 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 159,859 / month, #10,687 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaProgramming Language :: Python :: 3.9 |
Evidence: dataengine-0.0.92-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “spark sql workflow framework”
- dataengineDataengine is a Python framework for orchestrating data pipelines…
- brickflowsA Python framework and CLI tool for building and deploying workflows…
- fugueFugue provides a unified Python interface to write code once and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also pyspark · newtools · pyspark-client · dbt-databricks · policyengine-us · dagster-spark · pyspark-extension · koalas · delta-sharing · pyspark-pandas