dbl-discoverx
DiscoverX - Map and Search your Lakehouse
Decision gist · record as of 2026-08-14
Yes, if you are a Databricks user managing large Lakehouses and can tolerate the aging maintenance status and unclear license. The low install friction and no known vulnerabilities are positive signals. However, verify the license terms and test compatibility with your current Databricks and Python versions before production use, and understand that support is community-driven, not SLA-backed.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a Databricks workspace with Unity Catalog enabled; designed to run in Databricks notebooks, not standalone Python environments.
- Low install friction with a single runtime dependency (pyyaml).
- Package status is aging—last release was 469 days ago—and the project is provided without formal SLAs by Databricks Labs, meaning maintenance and support are exploratory only.
License · maintenance · safety
(unclear) — License treatment is unclear; the classifier indicates 'Other/Proprietary License' but no SPDX identifier or raw license text is available. Verify the actual license terms before adopting in production or proprietary work.
last release 2025-05-02 (469 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 541,881 downloads/mo, #6,092 on PyPI
Alternatives
Verify before relying
pip install dbl-discoverx
from discoverx import DX
dx = DX(locale="US")
result = dx.from_tables("catalog.schema.*").with_sql("OPTIMIZE {full_table_name}").apply()- Whether the package works with current Databricks SDK versions and recent Python releases
- Specific Python version requirements (requires_python is unspecified)
- Whether aging status reflects active maintenance or dormancy
What it is and what it does
DiscoverX is a Databricks Labs utility for automating administration across large numbers of Lakehouse assets. It lets you apply SQL templates or Python functions concurrently to multiple Delta tables selected by pattern matching, returning results as a unioned Spark DataFrame. Common use cases include bulk OPTIMIZE and VACUUM operations, deep cloning, tag management, owner changes, and semantic classification of columns for governance tasks like PII detection and GDPR compliance.
The package depends only on pyyaml and is designed to run in Databricks notebooks. It provides a fluent API (the `DX` class) for selecting tables, filtering by column presence, and executing templated operations with configurable concurrency. The fact sheet indicates the project is aging (469 days since last release) and carries no formal support SLA from Databricks—it is provided as-is for exploration.
Use it for
- Run OPTIMIZE and VACUUM across all tables in a schema or catalog to reclaim storage and improve query performance.
- Deep clone entire schemas or catalogs for testing or disaster recovery without manual table-by-table operations.
- Detect and report on tables with too many small files or ineffective ZORDER clustering across the entire lakehouse.
- Scan all tables for PII (email, phone, IP) using semantic classification rules and extract or delete sensitive data for GDPR compliance.
- Bulk update table ownership or apply tags across hundreds of assets matching a selection pattern.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are a Databricks user managing large Lakehouses and can tolerate the aging maintenance status and unclear license.
The low install friction and no known vulnerabilities are positive signals. However, verify the license terms and test compatibility with your current Databricks and Python versions before production use, and understand that support is community-driven, not SLA-backed.
Install
dbl-discoverx on PyPI
Before you install
Low install friction with a single runtime dependency (pyyaml). Package status is aging—last release was 469 days ago—and the project is provided without formal SLAs by Databricks Labs, meaning maintenance and support are exploratory only.
Requires a Databricks workspace with Unity Catalog enabled; designed to run in Databricks notebooks, not standalone Python environments.
License in practice
License treatment is unclear; the classifier indicates 'Other/Proprietary License' but no SPDX identifier or raw license text is available. Verify the actual license terms before adopting in production or proprietary work.
Quickstart
pip install dbl-discoverx
from discoverx import DX
dx = DX(locale="US")
result = dx.from_tables("catalog.schema.*").with_sql("OPTIMIZE {full_table_name}").apply()
Verify before relying
- Whether the package works with current Databricks SDK versions and recent Python releases
- Specific Python version requirements (requires_python is unspecified)
- Whether aging status reflects active maintenance or dormancy
Package facts
| License | Not declared unclear |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagepyyaml |
| Maintenance | Aging 469 days since the last release |
| First released | |
| Downloads | 541,881 / month, #6,092 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: Other/Proprietary LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: dbl_discoverx-0.0.9-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “lakehouse bulk operations”
- dbl-discoverxDiscoverX automates bulk administration tasks across Lakehouse assets…
- bauplanBauplan is a CLI and SDK for interacting with a code-first data…
- dbt-fabricsparkA dbt adapter that connects dbt to Apache Spark in Microsoft Fabric…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also dbt-databricks · dbldatagen · databricks-labs-lsql · deltalite · databricks-sql · databricks-dlt · featuretools · dlt · databricks-labs-dqx · databricks-feature-engineering