Packages
Provides a Python client library for managing Azure Batch resources—accounts, pools, applications, and jobs—through the Azure Resource Manager API.
Autobahn provides WebSocket and WAMP (Web Application Messaging Protocol) implementations for Python, enabling bidirectional real-time messaging and remote procedure calls over WebSocket on Twisted and asyncio.
Install it if you need WebSocket or WAMP for real-time messaging.
pywinrm is a Python client for Windows Remote Management (WinRM) that lets you run commands, PowerShell scripts, and fetch WMI data on remote Windows machines from any machine running Python.
Install it if you need to automate Windows systems from Python.
Runs Cloud Custodian policies in parallel across multiple AWS accounts, Azure subscriptions, GCP projects, or OCI tenancies from a single configuration file.
Install it if you use Cloud Custodian and manage more than one account or subscription.
Provides Python client library for managing Azure Service Fabric clusters, applications, and node types through the Azure Resource Manager API.
Install it if you need to manage Service Fabric resources from Python code.
Manages Azure Batch AI resources (clusters, jobs, experiments, workspaces) via the Azure Resource Manager API; deprecated as of 10-31-2024 in favor of azureml-core.
Provides enhanced HTTPS support for Python's httplib and urllib2 (or http.client and urllib in Python 3) by wrapping PyOpenSSL to enable full SSL peer verification using certificate validation.
HdfsCLI provides Python bindings and a command-line interface for interacting with HDFS clusters via the WebHDFS API, supporting file operations, metadata queries, and an interactive shell.
However, the package is dormant, so expect no active support or fixes for new Hadoop versions.
Kazoo provides a higher-level Python client API for Apache ZooKeeper, simplifying coordination, configuration management, and distributed locking operations.
Install it if you're building on ZooKeeper for coordination; skip it only if you don't need ZooKeeper or prefer a different coordination backend.
Integrates Dagster data pipeline orchestration with Kubernetes, enabling declarative asset management and orchestration on k8s infrastructure.
Install only if you have a running k8s cluster and need Dagster's orchestration on that infrastructure; otherwise, use Dagster's local or other executor options.
E2B provides a Python SDK to create and control isolated cloud sandboxes for executing untrusted or AI-generated code safely, with support for command execution and code interpretation.
Stamina wraps Tenacity to provide a production-ready retry decorator with sensible defaults: exponential backoff with jitter, attempt and timeout limits, exception filtering, and built-in async support for asyncio and Trio.
An asyncio Python client for connecting to and messaging through NATS, a distributed messaging system, with support for core pub-sub, request-reply, and JetStream persistence patterns.
Install it if you need lightweight async messaging, pub-sub distribution, or persistent queues in a distributed architecture.
RedBeat is a Celery Beat scheduler that persists scheduled tasks and runtime state in Redis instead of the local filesystem, enabling dynamic task management and multi-instance coordination.
Install it if you need dynamic task scheduling or multi-instance Beat coordination; skip it if you're using a single Beat instance with a static task set.
Dagster-aws provides AWS-specific integrations for Dagster, enabling data pipeline orchestration to work with AWS services like S3, EC2, and other AWS resources.
Provides a Python client library to manage Azure NetApp Files resources—volumes, capacity pools, accounts, snapshots, and backups—via the Azure Resource Manager API.
Install it if you need to manage Azure NetApp Files programmatically; do not install it if you only consume NetApp storage via other Azure services without needing…
Manages Azure Service Fabric Managed Clusters through a Python client library, providing operations for applications, services, node types, and cluster configuration.
Provides mutable mapping tools and data structures for working with dictionary-like interfaces in Python.
CLI tool for managing and deploying Dagster data pipelines to Dagster Cloud, providing command-line access to cloud orchestration and asset management features.
Install only if you have a Dagster Cloud instance to manage; it is not useful as a standalone tool for local-only Dagster development.
arq provides a job queue system for Python that uses asyncio and Redis to run background tasks asynchronously and distribute work across multiple workers.
Stores Celery task results in a Django database or cache backend, making task completion status and results queryable through the Django ORM.
Integrates Celery as a distributed task execution backend for Dagster data pipelines, enabling horizontal scaling of asset materialization and job runs across multiple worker nodes.
Install only if you have a Celery broker (RabbitMQ, Redis, etc.) already running or planned; it adds complexity that single-machine deployments do not need.
Integrates Apache Airflow DAGs with OpenLineage to automatically collect and emit task lineage, metadata, and data provenance events.
dagster-dg-core provides the core orchestration and asset-definition framework for Dagster, enabling you to declare data assets as Python functions and manage their execution and dependencies.
Install it if you are building data pipelines in Python and want declarative asset management with built-in lineage and observability.
Provides Python bindings to the Vast.ai API for searching GPU compute offers, managing instances, and accessing serverless inference endpoints.
Provides Python client bindings for the Kubeflow Pipelines API, enabling programmatic interaction with Kubeflow Pipelines services running on Kubernetes.
Not recommended if you're looking for a high-level SDK—use the main Kubeflow Pipelines SDK instead.
Provides a Python client to interact with Kestra servers for triggering flows, sending metrics and outputs, and retrieving execution status and logs.
Install it if you are already using or planning to use Kestra for workflow orchestration.
Integrates Dagster data orchestration with Docker, enabling containerized execution of data pipelines and assets within Dagster's orchestration framework.
Install it if you need to run Dagster assets in Docker containers.
Integrates dbt data transformation workflows into Dagster's orchestration platform, allowing you to define, schedule, and monitor dbt models as part of a larger data asset graph.
Install it if you are using Dagster for orchestration and want to include dbt models in your asset graph; skip it if you are managing dbt independently or not using…
Adds Spark to Python's import path at runtime, making it available as a regular library without manual sys.path manipulation or symlinking.
Install it if you work with Spark and want to avoid manual sys.path or environment variable management.
Runs AI-generated or arbitrary Python code in isolated cloud sandboxes with state preservation across multiple executions.
Python client for connecting to Apache Spark clusters via Spark Connect, enabling distributed data processing and analytics from Python without requiring local Spark JARs.
Python wrapper for Apache Sedona, a cluster computing system that extends Apache Spark with spatial data processing capabilities for loading, processing, and analyzing large-scale geographic data across distributed machines.
BullMQ for Python is a Redis-backed job queue library that lets you add, schedule, and process jobs with features like priority, deduplication, retries, and worker management.
Provides async Python bindings to the Vers REST API for managing virtual machines and orchestration through the Vers control plane.
However, adoption appears early-stage (zero repository stars, recent first release in 2026-04-17), so evaluate whether Vers itself is stable and suitable for your use…
Provides a CLI and Python SDK for managing GPU compute resources on Vast.ai, including instance lifecycle operations, resource searching, and serverless endpoint inference.
Cloud Custodian is a rules engine that enforces cloud infrastructure policies across AWS, Azure, and GCP by defining filters and actions in YAML to manage compliance, security, and cost.
Install it if you need to automate cloud compliance, security, or cost management across AWS, Azure, or GCP at scale.
Provides iPython magics for Amazon EMR Notebooks to mount S3 workspaces, generate presigned S3 download URLs, and execute notebooks in the background.
Integrates Dagster data orchestration with Celery task queue and Kubernetes execution, enabling distributed pipeline runs across containerized infrastructure.
dvc-task queues and runs background jobs from Python applications using Celery, without requiring a separate messaging server or broker infrastructure.