Subcategories
Packages
Parses SQL queries using sqlglot and extracts metadata: column names, table names, aliases, query types, and values from INSERT statements.
Install it if you need to extract or analyze SQL query structure—it handles the parsing complexity so you don't have to.
Provides a Rust-powered core implementation for Apache Iceberg table format operations in Python, intended as an internal dependency for the pyiceberg library.
Converts SQL queries between different parameter styles (named, ordinal, numeric, and dollar-sign variants) to match what your database driver supports.
Install it if you need to bridge parameter style mismatches between your code and your database driver.
Manages Azure Data Migration resources and orchestrates database migration tasks across SQL Server, MySQL, PostgreSQL, MongoDB, and Oracle sources to Azure targets.
dbt-spark is the Apache Spark adapter for dbt, enabling data transformation workflows in Spark using dbt's SQL and YAML-based modeling practices.
Install it if you use Apache Spark and want to adopt dbt's transformation and testing practices.
PyTables provides an object-oriented interface to HDF5 for storing and retrieving hierarchical datasets with efficient compression, designed to handle extremely large amounts of multidimensional data.
Docling Core defines the foundational DoclingDocument data model and provides APIs for serialization, chunking, and profiling of structured document data for generative AI applications.
dbt-bigquery is a dbt adapter that enables data transformation workflows in Google BigQuery, allowing analysts and engineers to organize, cleanse, and prepare raw warehouse data using dbt's SQL and YAML-based practices.
Install it if you use BigQuery and want to adopt dbt's SQL-based transformation practices; skip it if you prefer procedural data pipelines or are not yet using BigQuery.
Provides Python access to Snowflake entity metadata and resource management, allowing you to create, delete, and modify Snowflake resources programmatically.
Install it if you need to automate Snowflake object lifecycle operations from Python; skip it if you only need to query data (use snowflake-connector-python directly…
Python client library for connecting to and executing commands against a Valkey key-value store, with support for connection pooling, pipelining, pub/sub messaging, and both synchronous and asynchronous operations.
Native Python client library for connecting to and querying Vertica databases, implementing the DB-API v2.0 standard.
However, the aging maintenance status suggests it may not receive timely updates for new Vertica features or Python versions—verify compatibility with your specific…
Manages Azure PostgreSQL Flexible Server resources via the Azure Resource Manager API, allowing programmatic creation, configuration, and lifecycle operations on PostgreSQL database servers in Azure.
However, verify the license terms in the repository before use in proprietary projects, since the license treatment is currently unclear in the package metadata.
PyExasol is the official Python connector for Exasol databases, optimized for high-throughput data transfer using WebSocket protocol and supporting parallel data streams for multi-core scalability.
Install it if you work with Exasol and need Python connectivity; the parallel-stream architecture makes it the natural choice over generic ODBC drivers for this DBMS.
sqlite-vec adds vector search capabilities to SQLite, enabling similarity queries and vector operations directly within SQLite databases.
Executes SQL queries against Databricks through the Python SDK with minimal dependencies, returning results as iterators or single values without requiring a persistent connection.
However, verify the unclear license terms before committing to a production deployment, and confirm that the REST-based result fetching meets your performance…
Provides a Python client library for managing Azure MySQL Flexible Server resources through the Azure management API.
Manages SQL Server virtual machines in Microsoft Azure, providing a Python client to create, configure, and monitor SQL VM resources via the Azure Resource Manager API.
Install it if you need to automate SQL Server infrastructure on Azure from Python.
A Python client for connecting to and querying the Hive metastore via the Thrift protocol, enabling programmatic access to Hive metadata and partition information.
However, note that the package has not been updated since 2018—verify that the Thrift protocol implementation works with your target Hive metastore version before…
A SQL parser, transpiler, and optimizer that converts SQL between over 30 dialects with no runtime dependencies.
However, the package summary marks it as deprecated in favor of sqlglotc—verify whether that migration is necessary for your use case before committing.
Provides a Python binding to LMDB (Lightning Memory-Mapped Database), a fast key-value store based on memory-mapped files for efficient data access and storage.
Parses PostgreSQL SQL statements into an abstract syntax tree and provides tools to inspect and prettify the parsed structure.
Wraps cachetools cache implementations with disk persistence, allowing cached data to survive across program restarts while maintaining the same cache interface.
However, be aware that the package is aging (last release 242 days ago) and has a known issue on Windows with Python 3.13+; verify that issue doesn't affect your…
PynamoDB provides a Pythonic ORM-like interface to Amazon DynamoDB, simplifying table definition, querying, and data operations compared to the raw AWS API.
Install it if you're building on DynamoDB and want a cleaner API than raw boto3; skip it only if you need very low-level control or are already committed to a…
Provides base classes and utilities for building database testing frameworks, allowing you to create isolated, temporary database instances for test suites.
Automatically creates and tears down temporary PostgreSQL instances for testing, handling setup and cleanup without manual intervention.
DataFusion is a Python binding to Apache Arrow's in-memory query engine, enabling SQL and DataFrame-based queries against CSV, Parquet, and JSON data with built-in query optimization.
Install it if you need SQL query capabilities over Parquet/CSV/JSON without building a custom query engine or loading entire datasets into memory.
Provides an object-oriented Python API to connect to and query LDAP directory servers, wrapping OpenLDAP client libraries and offering utilities for LDIF processing, LDAP URLs, and schema handling.
MetricFlow compiles metric definitions into reusable SQL queries, handling multi-hop joins, complex metric types, and aggregations across different time granularities.
Pure Python DBAPI driver for MSSQL that implements the TDS protocol without requiring ADO or FreeTDS, enabling database connections on any platform.
Install it if you need to connect to MSSQL from Python without external system dependencies.
Extends PySTAC to describe tabular data assets with column definitions, row counts, keys, and storage formats via the Table Extension specification.
Extends PySTAC to support the Item Assets Definition Extension, enabling collections to define shared asset properties as templates across all items.
Generates Entity Relation (ER) diagrams from SQLAlchemy models or existing databases, outputting to image, PDF, or markdown formats.
Install it if you need to visualize database schemas or SQLAlchemy models; the GraphViz system dependency is the only real prerequisite.
dbt-redshift is an adapter that enables dbt to transform and manage data directly in Amazon Redshift warehouses using dbt's SQL and YAML-based workflow.
Install it if you use Amazon Redshift and want to adopt dbt's version-controlled, software-engineering approach to data transformation.
Writes and queries InfluxDB 3.0 using SQL or InfluxQL, with support for Point objects, line protocol, DataFrames, and file imports.
Unified Python API for Snowflake workloads, providing access to data engineering, Snowpark, Snowpark ML, and client application resources through a single namespace package.
PySide6 Essentials provides Python bindings to Qt 6.0+ core modules for building cross-platform desktop applications with native UI widgets, graphics, networking, and database support.
The copyleft license requires careful review for commercial projects; consider a commercial Qt license if needed.
cx_Oracle provides a Python interface to Oracle Database, implementing the Python database API 2.0 specification to enable queries, transactions, and data manipulation against Oracle instances.
However, the aging release cycle (last update 2021-11-04) and lack of explicit support for Python 3.11+ make it less ideal for new projects targeting modern Python…
Shiboken6 provides Python access to metadata and introspection functions for C++/Python bindings, allowing you to check object validity and debug wrapper state in applications built with Qt for Python.
Install only if you actually need binding introspection; it is not a general-purpose library.
Provides temporary backward compatibility for code that imports the old unrelated `snowflake` package, allowing it to read from `/etc/snowflake` or an alternative path; this is a migration bridge, not a primary tool.
PySide6-Addons extends PySide6 with additional Qt modules for 3D graphics, multimedia, web engines, charts, and specialized hardware interfaces.