Packages
Integrates Databricks AI features, particularly vector search, into OpenAI applications through tool definitions and retrieval workflows.
The package is narrowly scoped to a specific integration pattern, so install only if you have both Databricks infrastructure and OpenAI API access.
Provides Python bindings to all public Databricks REST APIs, enabling programmatic access to workspace resources, clusters, jobs, and other Lakehouse operations.
Install it if you need programmatic access to Databricks workspaces from Python.
Provides a Python interface for querying and manipulating Databricks SQL tables with chainable methods for select, insert, update, and delete operations.
A Python client library that connects to Databricks clusters and SQL warehouses using a Thrift-based protocol, conforming to the Python DB API 2.0 specification and supporting Arrow-based data exchange.
Install it if you need to query Databricks clusters or SQL warehouses from Python.
Provides a SQLAlchemy 2.0 dialect that bridges SQLAlchemy applications to Databricks SQL, enabling ORM and SQL expression language support for Databricks workspaces.
Install it if you need to use SQLAlchemy with Databricks.
Provides a unit testing framework for Databricks notebooks, allowing you to run notebook code locally with mocked Databricks objects like spark, dbutils, and display for pytest-based testing.
A deprecated re-export shim that forwards to databricks-ai-search and emits a deprecation warning; existing code using databricks.vector_search.* imports continues to work with backward-compatible aliases.
A Python client for streaming data ingestion into Databricks Delta tables via the Zerobus service, supporting both JSON and Protocol Buffer serialization with synchronous and asynchronous APIs.
A Python wrapper for the Databricks REST API that provides programmatic access to Databricks workspace resources including tokens, secrets, clusters, jobs, and DBFS.
Reads and writes CSV files using Python dataclasses, with automatic type conversion, validation, and error reporting tied to specific CSV line numbers.
Install it if you work with CSV files and want type safety and cleaner code than dict-based approaches; skip it only if you need advanced CSV features (e.g.,…
Converts dataclasses to and from dictionaries and other common types with minimal configuration, supporting nested structures, enums, unions, and custom validation.
Dataclass Wizard converts Python dataclasses to and from JSON, YAML, TOML, and environment variables with automatic type coercion and case transformation, using code-generated serialization for performance.
Install it if you need JSON/YAML/TOML/environment variable marshalling for dataclasses and want automatic type coercion and case transformation out of the box.
Provides the PEP 557 dataclasses module for Python 3.6, enabling decorator-based class definitions with automatic generation of `__init__`, `__repr__`, and other standard methods.
Generates Avro schemas from Python dataclasses, Pydantic models, and Faust Records, and serializes/deserializes Python instances to and from Avro binary or JSON formats.
Install it if you need to work with Avro schemas in Python—either to generate them from your type definitions or to serialize/deserialize Avro data.
Converts Python dataclasses to and from JSON with minimal boilerplate, supporting nested structures, custom field naming, and schema validation.
Install it if you work with dataclasses and JSON; the decorator is simple to apply and the API is straightforward.
Converts Python dataclasses to and from JSON with a decorator or mixin, supporting nested types, collections, datetime, UUID, and Decimal objects.
Install it if you need straightforward JSON serialization for dataclasses; the two-dependency footprint and decorator-based API make it a lightweight choice.
Generates JSON Schema from Python dataclasses and provides serialization, deserialization, and validation against the generated schema.
DataComPy compares two DataFrames across Pandas, Polars, Spark, and Snowflake, reporting differences in rows and columns with configurable matching tolerance and structured output for programmatic access.
Install it if you need to compare tabular data programmatically or generate detailed difference reports.
A CLI tool for defining, validating, and testing data contracts using the Open Data Contract Standard, with support for schema validation, quality checks, and exports across multiple data platforms.
Reads and writes YAML files conforming to the Data Contract Specification using Pydantic models, enabling programmatic access to data contract definitions.
However, the aging maintenance status (323 days since last release) suggests checking whether the package is still actively developed before adopting it for critical…
Datacube provides an integrated gridded data analysis environment for managing and analyzing decades of analysis-ready earth observation satellite data from multiple acquisition systems.
Generates human-readable diffs of Python data structures (dicts, lists, tuples, sets, strings) and provides drop-in assertion replacements for test frameworks.
However, avoid it if you need ongoing support or compatibility assurance with current Python versions.
Provides a Python client library to send metrics, events, and service checks to Datadog via its HTTP API and DogStatsD protocol, with support for both UDP and Unix domain socket transports.
Install it if you are already using Datadog and need to send metrics or events from Python; if you need comprehensive API coverage for all Datadog endpoints, consult…
Python client library for the Datadog API, providing programmatic access to Datadog's monitoring, alerting, dashboards, and incident management endpoints.
Install it if you need programmatic access to Datadog's monitoring, alerting, dashboards, or incident management from Python.
Provides AWS CDK constructs to automatically instrument Lambda functions and ECS Fargate services with Datadog monitoring, configuring metrics, traces, and logs collection without manual layer management.
Install it if you are already using AWS CDK v2 and want Datadog observability; it saves configuration effort and enforces consistent instrumentation across your…
Provides the Python base classes and utilities that Datadog Agent integrations need to run, and mocks an Agent environment for standalone testing and development.
Install it if you are building, testing, or maintaining Datadog checks.
Sends Python log messages to Datadog as Events in the Events Explorer via a custom logging.Handler, with support for tags and mentions.
Dataengine is a Python framework for orchestrating data pipelines that integrates pandas, Apache Spark, and cloud services (AWS, Databricks, GitHub, Slack, Datadog) through a configuration-driven Engine class that manages Database, Dataset, and Query objects.
However, the Alpha status, 484-day staleness, unclear license, and heavy dependency footprint (18 runtime deps) mean you should verify the license terms, confirm the…
Datafiles is a file-based ORM that automatically synchronizes Python dataclasses to disk files (YAML, JSON, TOML, JSON5) and back, treating the filesystem as a bidirectional persistence layer with minimal boilerplate.
Install it if you need automatic file-based persistence for configuration, state, or fixtures; skip it if you require a traditional database or need Python versions…
Reads and writes tabular data in multiple formats (CSV, XLS, JSON, ODS, and others) from local files, HTTP, FTP, and S3 sources with low memory overhead.
However, be aware that the last release was 873 days ago and no active development is occurring—security patches and bug fixes are unlikely, so audit the dependencies…
A Python client library for the DataForSEO REST API, handling request construction, response parsing, and authentication for SEO data retrieval across 12 API sections including SERP, keywords, domain analytics, backlinks, and more.
However, resolve the unclear license status before use in commercial or copyleft-sensitive contexts, and ensure you have valid DataForSEO API credentials and…
Provides a lightweight compatibility layer that lets you write dataframe code once and run it against pandas, Polars, or any other library implementing the Python Dataframe API Standard.
However, maintenance is dormant (no commits since April 2024), so verify that the standard implementation covers your actual use cases and that the libraries you…
Dataframely validates the schema and content of Polars data frames using declarative class-based schemas with field constraints and custom validation rules.
However, verify the license before use in proprietary work, and be aware that the package is relatively new (first release March 2025)—test it in non-critical…
DataFusion is a Python binding to Apache Arrow's in-memory query engine, enabling SQL and DataFrame-based queries against CSV, Parquet, and JSON data with built-in query optimization.
Install it if you need SQL query capabilities over Parquet/CSV/JSON without building a custom query engine or loading entire datasets into memory.
Builds DataFusion SQL queries programmatically with injection-safe value handling, emitting readable SQL text rather than requiring string concatenation or templates.
This is a redirect package that installs acryl-datahub when you run `pip install datahub`, preventing name squatting and ensuring users get the correct data governance tool.
A Python SDK for calling the Datalab API to convert documents to markdown and execute multi-step document processing workflows.
Generates Python data models (Pydantic, dataclasses, TypedDict, msgspec) from schema definitions like OpenAPI, JSON Schema, Protocol Buffers, GraphQL, and raw data formats.
Datamol provides a pythonic layer on top of RDKit for molecular manipulation, offering simplified APIs for converting between molecular formats, standardizing molecules, and performing common cheminformatics operations.
Install it if you work with molecular structures and want a more ergonomic API than raw RDKit; the Apache-2.0 license poses no barrier.
Provides classes and utilities to create, load, validate, and manipulate Data Packages—standardized containers for tabular and non-tabular datasets with machine-readable metadata.
However, dormant maintenance (885 days since last release) and outdated Python version classifiers (2.7–3.7) are concerns for new projects; evaluate the Frictionless…