Packages
oneDAL is a C++ and DPC++ library that implements accelerated machine learning routines for tabular data (linear regression, K-means clustering, random forests, etc.) for CPUs, GPUs, and distributed setups, with a Python interface powered by tbb.
However, the unclear Intel Simplified Software License requires verification before commercial use, and installation is limited to x86_64 Windows or manylinux_2_28…
daal4py provides a Python API to Intel's oneAPI Data Analytics Library, offering optimized implementations of machine learning and data analytics algorithms.
Not recommended if you rely on scikit-learn patching—use Intel Extension for Scikit-learn instead.
Converts dictionaries into dataclass instances with support for nested structures, type checking, optional fields, unions, generics, and custom type hooks.
Dacktool provides unspecified Python utilities with no runtime dependencies.
Install only if you are maintaining legacy code that already depends on it.
Daemonize provides a Python library for converting regular Python scripts into Unix daemon processes that run in the background with proper process management.
However, do not use it for new projects or for production systems where you need active maintenance and security updates.
Compares two tables and produces a diff summary that can be used as a patch file, optimized for tables sharing a common origin.
Provides the Python runtime library needed to execute Dafny-compiled Python code, enabling verified programs written in Dafny to run on Python.
Daft is a distributed dataframe engine for processing images, audio, video, and structured data at scale, with built-in AI operations and support for multimodal workloads.
dag-factory builds Apache Airflow DAGs declaratively from YAML configuration files, eliminating the need to write Python code for DAG construction.
Install it if you manage multiple similar Airflow workflows or want to let non-Python developers define DAGs; skip it only if your DAGs are highly custom or require…
Dagger Python SDK is a client for defining and running CI/CD pipelines in Python, executing them on any OCI-compatible container runtime without requiring manual container orchestration.
Install it if you want to replace shell scripts and YAML with type-safe, testable Python—but only if you have Docker or another OCI runtime available and are…
Dagit is the web UI for Dagster, providing a visual interface to orchestrate, monitor, and manage data pipelines and assets defined as Python functions.
Install it if you are already using Dagster or evaluating it for data pipeline orchestration; it is essential for visual monitoring and management of Dagster assets.
Dagster is a data pipeline orchestrator that lets you declare data assets as Python functions and automatically runs them at the right time to keep those assets up-to-date, with built-in lineage tracking and observability.
Install it if you need declarative asset-based orchestration with integrated lineage and observability; skip it if you prefer lightweight task scheduling or are…
Integrates Airbyte data connectors with Dagster's orchestration engine to define and run data ingestion assets as part of declarative data pipelines.
Install it if you are already using Dagster for orchestration and want to integrate Airbyte connectors as declarative assets; it is not necessary if you are not using…
Dagster-aws provides AWS-specific integrations for Dagster, enabling data pipeline orchestration to work with AWS services like S3, EC2, and other AWS resources.
Provides Azure-specific integrations for Dagster, enabling data pipelines to interact with Azure services like Blob Storage, Data Lake, and Machine Learning.
Integrates Celery as a distributed task execution backend for Dagster data pipelines, enabling horizontal scaling of asset materialization and job runs across multiple worker nodes.
Install only if you have a Celery broker (RabbitMQ, Redis, etc.) already running or planned; it adds complexity that single-machine deployments do not need.
Integrates Dagster data orchestration with Celery task queue and Kubernetes execution, enabling distributed pipeline runs across containerized infrastructure.
Dagster Cloud is a managed orchestration platform that deploys and runs data pipelines with serverless or hybrid infrastructure, providing CLI tooling and agent management for Dagster workflows.
However, verify first whether your use case requires a Dagster Cloud account and whether the feature set (branching, CI/CD) you need is included in your tier.
CLI tool for managing and deploying Dagster data pipelines to Dagster Cloud, providing command-line access to cloud orchestration and asset management features.
Install only if you have a Dagster Cloud instance to manage; it is not useful as a standalone tool for local-only Dagster development.
Integrates Databricks with Dagster's data orchestration framework, enabling you to define and run data pipelines that interact with Databricks clusters and SQL warehouses.
Install only if you need Databricks-specific orchestration components; the base dagster library alone may suffice for simpler use cases.
Integrates Datadog monitoring and observability with Dagster data pipelines, enabling metric collection, alerting, and performance tracking within Dagster's orchestration framework.
Integrates dbt data transformation workflows into Dagster's orchestration platform, allowing you to define, schedule, and monitor dbt models as part of a larger data asset graph.
Install it if you are using Dagster for orchestration and want to include dbt models in your asset graph; skip it if you are managing dbt independently or not using…
A command-line tool for developing and managing Dagster data pipelines locally, providing scaffolding, code generation, and project initialization for Dagster-based data orchestration.
dagster-dg-core provides the core orchestration and asset-definition framework for Dagster, enabling you to declare data assets as Python functions and manage their execution and dependencies.
Install it if you are building data pipelines in Python and want declarative asset management with built-in lineage and observability.
Integrates dlt data loading into Dagster pipelines, enabling declarative ETL/ELT workflows where dlt sources and destinations are orchestrated as Dagster assets.
Install it if you are already using both Dagster and dlt and want to orchestrate dlt jobs as first-class Dagster assets rather than external processes.
Integrates Dagster data orchestration with Docker, enabling containerized execution of data pipelines and assets within Dagster's orchestration framework.
Install it if you need to run Dagster assets in Docker containers.
Integrates DuckDB with Dagster's data pipeline orchestration, providing DuckDB-specific ops and resources for declaring and running data assets.
Install it if you are already using Dagster and want to work with DuckDB; it adds minimal overhead and is backed by a well-resourced open-source project.
Provides ETL/ELT integration for Dagster, enabling data pipeline orchestration with embedded extraction and loading capabilities through dagster-dlt and dagster-sling connectors.
Integrates Fivetran data connectors with Dagster's orchestration engine, enabling declarative asset definitions that automatically sync data from Fivetran sources.
Provides GCP-specific integrations for Dagster, enabling data pipelines to interact with Google Cloud services like BigQuery and Cloud Storage.
Provides a GraphQL API interface for querying and interacting with Dagster data pipelines, assets, and orchestration metadata.
Integrates Dagster data pipeline orchestration with Kubernetes, enabling declarative asset management and orchestration on k8s infrastructure.
Install only if you have a running k8s cluster and need Dagster's orchestration on that infrastructure; otherwise, use Dagster's local or other executor options.
Integrates MLflow experiment tracking and model management with Dagster's data orchestration, allowing you to track ML model training and artifacts as part of your Dagster asset pipelines.
A Microsoft Teams integration resource for Dagster that enables posting messages and notifications to Teams channels from data pipeline orchestration workflows.
Integrates PagerDuty incident management with Dagster data pipelines, enabling automated alerting and incident response when pipeline events occur.
Integrates Pandera data validation with Dagster's asset orchestration, enabling schema validation checks within data pipeline workflows.
Dagster-pipes provides a toolkit for running Dagster integrations and transform logic outside of the main Dagster process, enabling external execution of data pipelines.
Connects Dagster data pipelines to PostgreSQL for asset storage, state management, and run tracking.
Install only if you have a PostgreSQL instance available and require multi-user or persistent storage; Dagster's default storage is sufficient for local development.
Integrates Prometheus metrics collection and export into Dagster data pipelines, enabling monitoring and observability of pipeline runs and asset computations.
Install only if you have a Prometheus deployment in place or plan to set one up; otherwise, Dagster's built-in observability may suffice.
Integrates PySpark with Dagster's data pipeline orchestration, enabling you to define and run Spark-based data assets within Dagster's declarative asset model.