Packages
Dask-awkward integrates Awkward Array with Dask to enable distributed, lazy computation on complex, nested data structures across multiple cores or machines.
Install it if you work with Awkward Arrays at scale or need lazy evaluation of nested structures across multiple workers.
Dask CUDA provides utilities for deploying and managing Dask workers on multi-GPU systems, enabling distributed computation across CUDA-enabled hardware.
Dask cuDF extends Dask DataFrame with a GPU-accelerated backend, enabling parallel and larger-than-memory data processing on NVIDIA GPUs using a pandas-like API.
Dask Expressions provides query optimization for Dask DataFrames by encoding operations as an expression tree that is optimized before execution, replacing the earlier Dask DataFrame implementation.
Install it if you are using Dask DataFrames; it is the recommended path forward.
Dask-GeoPandas parallelizes geospatial operations by combining GeoPandas' spatial capabilities with Dask's distributed computing, allowing you to process large geographic datasets across multiple partitions.
Install it if you work with large geographic datasets and need parallel processing.
Distributed generalized linear model fitting using Dask for parallel computation across clusters or multi-core systems.
Distributed image processing using Dask, enabling large-scale image operations across multiple cores or machines by parallelizing computations that would otherwise be memory-bound on a single system.
However, Pre-Alpha status means the API may change and feature coverage is incomplete—evaluate whether the available operations match your use case before committing…
Deploys Dask distributed computing clusters on job queuing systems like PBS, Slurm, SGE, and LSF, automating the submission and management of worker nodes.
Dask-ML provides distributed and parallel machine learning by integrating Dask with scikit-learn, XGBoost, and other ML libraries to scale training and inference across clusters.
However, note the aging maintenance status—last release was February 2025—so verify that the version meets your stability and security requirements before adopting in…
Generates high-quality synthetic datasets from scratch or seed data, with control over field relationships, statistical distributions, and built-in validation and quality scoring.
Provides a configuration API for building synthetic data generation pipelines in the NeMo Data Designer framework, allowing you to define data sources, LLM models, and column generation rules.
Execution engine for the NeMo Data Designer synthetic data generation framework, handling data transformation, generation, and LLM integration for creating synthetic datasets.
Provides version checking and validation utilities for distributed data platform components, plus data interface abstractions for charm libraries.
Provides PEP-561-compliant type stubs for NumPy, pandas, and Matplotlib, enabling mypy and other type checkers to recognize and validate types in these libraries.
Converts Python dictionaries and lists into XML documents with configurable structure and formatting.
Install it if you need dict-to-XML conversion and can work with Python 3.11+.
Provides async database access for PostgreSQL, MySQL, and SQLite using SQLAlchemy Core expressions, designed for integration with async web frameworks.
Python client library for connecting to Databend databases, supporting both local embedded in-process analytics and remote server connections with synchronous, asynchronous, and PEP 249 cursor interfaces.
Official Python client for accessing live and historical market data from Databento, supporting multiple asset classes, schemas, and data formats with normalized message structures.
Install it if you need programmatic access to Databento's historical or live market data; skip it if you don't have a Databento account or need data from a different…
databento-dbn provides Python bindings for encoding and decoding Databento Binary Encoding (DBN), a binary format for financial market data.
Databind deserializes JSON-like nested data structures into Python dataclasses and native types, and serializes dataclasses back to JSON-compatible dicts.
Not recommended if you prioritize serialization speed—the docs explicitly point to mashumaro for high-performance use cases.
Deserializes and serializes Python dataclasses and native types to and from JSON-like nested data structures, with support for enums, decimals, UUIDs, paths, datetimes, and generic types.
Deserializes and serializes Python dataclasses to and from JSON-like nested data structures, supporting native Python types, enums, dates, UUIDs, and generic types with customizable serialization behavior.
Not recommended if performance is critical—the maintainers explicitly direct high-speed use cases elsewhere.
Mosaic AI Agent Framework SDK provides a Python library for building and deploying agents on the Databricks platform, integrating with Databricks services and LLM capabilities.
Provides a shared Python API layer for building agents and applications that integrate with Databricks AI features like Vector Search and AI/BI Genie.
Python client for Databricks AI Search, a managed vector and keyword search service on the Databricks platform, supporting endpoint creation, index management, and similarity search queries.
However, the license is restrictive and proprietary—use is only permitted in connection with a Databricks Platform Services agreement.
Provides a simplified Python interface to the Databricks REST API by wrapping the databricks-cli client library, exposing service instances for jobs, clusters, policies, workspaces, and other Databricks resources.
Provides runtime support for Databricks AutoML, enabling automated machine learning model training and evaluation workflows within the Databricks platform.
Extends Databricks Declarative Automation Bundles to define jobs and pipelines as Python code, dynamically generate them from metadata, and modify bundle definitions during deployment.
A command-line interface for interacting with Databricks REST APIs, now superseded by newer versions and a dedicated SDK.
Install only if you are maintaining legacy scripts that cannot be migrated; for any new integration, use the recommended CLI 0.200+ or databricks-sdk-py instead.
Databricks Connect is a client library that lets you write Spark code locally in your IDE or notebook and execute it remotely on a Databricks cluster instead of running it locally.
Provides a DBAPI 2.0 connection interface and SQLAlchemy dialects to query Databricks Workspace and SQL Analytics clusters using either pyhive or pyodbc backends.
No, not recommended for new projects.
Provides type hints, API specs, and IDE autocomplete support for developing Databricks Delta Live Tables pipelines locally without functional runtime implementation.
However, maintenance is dormant, so type hints may lag behind the current DLT API.
Databricks Feature Engineering client for creating, managing, and serving feature tables within Databricks workspaces, including training models on feature data and publishing to online stores.
However, the proprietary license restricts use to Databricks Platform Services; confirm your Databricks agreement permits this library before production deployment.
This package is deprecated as of v0.17.0; it provides a client for Databricks Feature Store but has been replaced by databricks-feature-engineering, which maintains backward compatibility with existing imports.
Provides Python-native pathlib-like interfaces for Databricks Workspace paths, plus TUI primitives, logging, parallel task execution, and application state management.
However, the license treatment is unclear—verify the actual license terms in the repository before adopting it in proprietary or commercial work.
DQX provides rule-based data quality checking for PySpark DataFrames on Databricks, supporting batch and streaming workloads with built-in checks, custom rules, and automated quality monitoring.
However, the license treatment is unclear—verify the actual license before production use.
Executes SQL queries against Databricks through the Python SDK with minimal dependencies, returning results as iterators or single values without requiring a persistent connection.
However, verify the unclear license terms before committing to a production deployment, and confirm that the REST-based result fetching meets your performance…
Converts SQL code between different database dialects and reconciles data during migration to Databricks from enterprise data warehouses and other ETL sources.
However, the unclear license and Alpha status mean you should verify licensing compliance and test thoroughly in a non-production environment first.
Integrates Databricks AI services—LLMs, vector search, embeddings, and Genie—into LangChain applications through a unified package.
Install only if you have a Databricks workspace and LangChain is your chosen framework; it adds no value as a standalone tool.
Provides helpers and utilities to integrate MCP (Model Context Protocol) servers with Databricks environments, including OAuth authentication support across notebooks, model serving, and local development.