Subcategories
Packages
SQLTrie implements a prefix tree (trie) data structure backed by SQL, combining trie search capabilities with persistent database storage.
dbt-duckdb connects dbt (a SQL/Python transformation framework) to DuckDB, an embedded OLAP database that can read and write CSV, JSON, and Parquet files directly without loading them first.
Packages the MongoDB Curator tool as a Python library, allowing you to invoke Curator operations directly from Python code or retrieve its binary path.
django-reversion adds version control to Django model instances, allowing you to roll back changes, recover deleted records, and track history through a simple admin interface.
Install it if you need model versioning, audit trails, or recovery capabilities in a Django project.
Reads an existing database schema and generates SQLAlchemy model code automatically, supporting multiple output formats including declarative classes, dataclasses, and SQLModel.
Kedro-Datasets provides data connectors for Kedro's DataCatalog, implementing AbstractDataset for formats like CSV, Excel, Parquet, JSON, SQL, and Spark DataFrames across local, network, and cloud storage.
ConnectorX loads data from databases directly into Python dataframes (Pandas, PyArrow, Polars, Modin, Dask) using a Rust backend optimized for speed and memory efficiency, with optional parallel loading via partitioning.
Install it if you regularly load data from databases; the one-line API and optional parallelism justify the medium wheel size.
Provides mutation-tracked JSON column types for SQLAlchemy that detect and persist changes to JSON data at the top level or nested within objects and arrays.
However, maintenance is dormant—last release was over two years ago—so verify compatibility with your specific SQLAlchemy and Python versions before adopting it in…
databento-dbn provides Python bindings for encoding and decoding Databento Binary Encoding (DBN), a binary format for financial market data.
Provides a DBAPI 2.0-compatible Python interface to PostgreSQL via the ADBC driver manager, enabling Arrow-native data exchange with PostgreSQL databases.
DBOS adds durable workflows and queues to Python applications by checkpointing execution state in Postgres, allowing programs to automatically recover from failures without external orchestration infrastructure.
Bridges PyCasbin access-control policies with SQLAlchemy-supported databases, enabling policy persistence and retrieval across PostgreSQL, MySQL, SQLite, Oracle, SQL Server, Firebird, and Sybase.
It is a straightforward adapter with a narrow, well-defined scope—use it when pycasbin and SQLAlchemy are your architecture; do not install it if you are not already…
Milvus Lite is a pure-Python local vector database that provides dense and sparse vector search, BM25 full-text search, and scalar filtering through a Milvus-compatible API, storing data in a local `.db` file or embedded gRPC server.
Gdbmongo provides GDB pretty printers and commands for inspecting MongoDB Server internals in live processes and core dumps, with support for multiple MongoDB versions.
SQLLineage parses SQL statements to extract and visualize data lineage—identifying source tables, target tables, and intermediate tables involved in data transformations.
Tentaclio opens and manages streams across multiple protocols (file, FTP, SFTP, S3, HTTP/HTTPS) and database connections using a unified URL-based interface, with automatic credential injection and pandas integration.
Beanie is an asynchronous Python object-document mapper for MongoDB that lets you define data models using Pydantic and interact with MongoDB collections through a Document-based API.
Provides S3 storage integration for tentaclio by bundling tentaclio and boto3 as a single installable package.
aioodbc provides async/await access to ODBC databases by wrapping pyodbc with asyncio support, using threads internally to avoid blocking the event loop.
However, if you require active maintenance or are starting a new greenfield project, consider whether a more actively maintained async driver for your specific…
Converts SQL code between different database dialects and reconciles data during migration to Databricks from enterprise data warehouses and other ETL sources.
However, the unclear license and Alpha status mean you should verify licensing compliance and test thoroughly in a non-production environment first.
Extends Alembic to autogenerate database migrations for PostgreSQL functions, views, materialized views, triggers, and policies alongside standard SQLAlchemy model changes.
Install it if PostgreSQL-specific schema objects are part of your migration strategy.
Provides type stubs and IDE autocompletion for aiobotocore's RDS service, enabling static type checking with mypy, pyright, and IDE code completion.
Install it if you use aiobotocore for RDS and want IDE support or static type checking—it adds no runtime overhead.
Provides LangChain abstractions for vector storage, chat history, and document management backed by PostgreSQL with support for both synchronous and asynchronous operations.
Install it if you are already using LangChain and need a Postgres-backed vector store or session manager; skip it if you do not use LangChain or prefer a different…
django-pgtrigger lets you define PostgreSQL triggers declaratively on Django models, enabling database-level enforcement of constraints, state transitions, and data operations without application code.
A CLI tool for defining, validating, and testing data contracts using the Open Data Contract Standard, with support for schema validation, quality checks, and exports across multiple data platforms.
Analyzes SQL statements to extract source and target tables, intermediate tables, and column-level lineage, using pluggable parsers (sqlfluff or sqlparse) and networkx for graph representation.
Automate creation, reading, and modification of Tableau extract (.hyper) files programmatically, enabling custom ETL workflows and data source integration outside Tableau's native connectors.
Install only if you're already committed to Tableau as your BI platform and have a specific ETL or data-source automation need.
django-auditlog logs changes to Django model instances, recording what changed, when, and which user made the change, storing the summary in JSON format.
Install it if you need audit logging for Django models and want a simple, proven solution.
Provides Postgres advisory locks, table locks, and lock management utilities for Django applications to prevent concurrent task execution and handle blocking locks.
MySQL driver written in Python that enables database connections and queries to MySQL servers.
CLI tool and Python library for creating, querying, and transforming SQLite databases from JSON, CSV, or TSV data, with built-in full-text search and schema migration support.
Provides a DBAPI 2.0-compatible Python interface to SQLite via the ADBC (Arrow Database Connectivity) driver, enabling standard SQL queries with Arrow table results.
This package is deprecated and should not be installed. It provided a Python client for connecting to a vector database service.
Reads and writes YAML files conforming to the Open Data Contract Standard using Pydantic models, enabling programmatic access to data contract specifications.
Intake provides a declarative data catalog system for describing, discovering, and loading datasets from multiple sources and formats, with support for remote storage and compute platforms.
However, note the active security advisory (GHSA-37g4-qqqv-7m99) and verify it does not affect your use case before deploying to production.
Exposes Qdrant vector search as a Model Context Protocol server, allowing LLM applications to store and retrieve semantic memories from a vector database via standardized MCP tools.
Implements keyset-based paging for SQLAlchemy queries, replacing offset-based pagination to avoid performance degradation and result skipping on large datasets.
PyMongoCrypt provides Python bindings for libmongocrypt, enabling client-side encryption for MongoDB drivers through cffi and cryptography.
Registers custom SQLite functions for ranking and analyzing FTS4 full-text search results, including BM25 scoring, TF-IDF ranking, and matchinfo decoding.
However, the package is abandoned (last update 2022-07-30) and has no active maintenance, so expect no updates for bugs or compatibility issues with future Python…
A dbt adapter that connects dbt to Apache Spark in Microsoft Fabric via Livy endpoints, enabling SQL-based data transformation workflows on Fabric Lakehouses with or without schema support.
Install it if you use dbt with Microsoft Fabric Spark and need to transform data in Lakehouses.