databricks-connect
Databricks Connect Client
Decision gist · record as of 2026-08-14
Yes, if you have a Databricks workspace and need to run Spark code from a local IDE or notebook. The low install friction, active maintenance, and zero known vulnerabilities make it safe to adopt. The Python 3.12 requirement and proprietary license are the main constraints—verify the license terms for your use case and ensure your environment supports Python 3.12.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.12 exactly; connection to a Databricks workspace with valid credentials configured (typically via DATABRICKS_HOST and DATABRICKS_TOKEN environment variables or Databricks CLI config).
- Low install friction with a pure-Python wheel.
- Active maintenance with a release 14 days ago.
License · maintenance · safety
Databricks Proprietary License (unclear) — Licensed under Databricks Proprietary License with unclear treatment. You should review the license terms directly before adopting in a commercial or redistributable context.
last release 2026-07-31 (14 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 12,249,777 downloads/mo, #1,332 on PyPI
Alternatives
Verify before relying
pip install databricks-connect==19.0.0
from databricks.connect import DatabricksSession
spark = DatabricksSession.builder.build()
df = spark.sql('SELECT * FROM my_table')
df.show()- Whether the proprietary license permits use in closed-source commercial applications or only in open-source contexts.
- Whether Python 3.12 requirement will be relaxed in future releases or remains pinned.
- Performance characteristics and latency overhead of remote Spark execution compared to local Spark.
What it is and what it does
Databricks Connect bridges your local development environment—IDE, notebook server, or custom application—to a remote Databricks cluster. Instead of installing and running Spark locally, you write code using familiar Spark APIs and execute it on a managed Databricks cluster, letting the cluster handle computation while your local machine remains lightweight.
The package depends on py4j for JVM interop, grpcio for remote communication, and standard data libraries like pandas, pyarrow, and numpy. It's actively maintained and has low install friction as a pure-Python wheel, but it locks you to Python 3.12 and requires a Databricks workspace with valid credentials to function.
Use it for
- Develop and test Spark jobs in your IDE without installing a local Spark cluster.
- Run interactive Spark queries from a notebook server connected to a shared Databricks workspace.
- Integrate Spark computation into a Python application running on a lightweight client machine.
- Prototype data pipelines locally before deploying them to production Databricks clusters.
- Use Databricks-managed infrastructure for compute-heavy workloads while keeping your development loop fast.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have a Databricks workspace and need to run Spark code from a local IDE or notebook.
The low install friction, active maintenance, and zero known vulnerabilities make it safe to adopt. The Python 3.12 requirement and proprietary license are the main constraints—verify the license terms for your use case and ensure your environment supports Python 3.12.
Install
databricks-connect on PyPI
Before you install
Low install friction with a pure-Python wheel. Active maintenance with a release 14 days ago. Requires Python 3.12 specifically, which narrows compatibility to current-version environments only.
Requires Python 3.12 exactly; connection to a Databricks workspace with valid credentials configured (typically via DATABRICKS_HOST and DATABRICKS_TOKEN environment variables or Databricks CLI config).
License in practice
Licensed under Databricks Proprietary License with unclear treatment. You should review the license terms directly before adopting in a commercial or redistributable context.
Quickstart
pip install databricks-connect==19.0.0
from databricks.connect import DatabricksSession
spark = DatabricksSession.builder.build()
df = spark.sql('SELECT * FROM my_table')
df.show()
Verify before relying
- Whether the proprietary license permits use in closed-source commercial applications or only in open-source contexts.
- Whether Python 3.12 requirement will be relaxed in future releases or remains pinned.
- Performance characteristics and latency overhead of remote Spark execution compared to local Spark.
Package facts
| License | Databricks Proprietary License unclear |
| Python support | Capped below the current Python release ==3.12.* |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 12 packagespy4jsixpandaspyarrowgrpciogrpcio-statusgoogleapis-common-protoszstandardnumpydatabricks-sdkpackagingsetuptools |
| Maintenance | Actively maintained 14 days since the last release |
| First released | |
| Downloads | 12,249,777 / month, #1,332 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableLicense :: Other/Proprietary LicenseProgramming Language :: Python :: 3.12Programming Language :: Python :: Implementation :: CPythonTyping :: Typed |
Evidence: databricks_connect-19.0.0-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “databricks cluster client”
- databricks-connectDatabricks Connect is a client library that lets you write Spark code…
- databricks-apiProvides a simplified Python interface to the Databricks REST API by…
- databricksapiA Python wrapper for the Databricks REST API that provides…
Give your agent the search over MCP, or paste the wish link into any chat.
More Distributed Computing packages
gRPC Python is an HTTP/2-based RPC framework that enables you to define and call remote procedures across network boundaries using protocol buffers for serialization.
Install it if you need RPC communication in a distributed system or are integrating with existing gRPC services.
execnet lets you spawn and communicate with Python interpreters across local processes, remote hosts, and different platforms, using a simple API for task distribution and inter-process messaging.
However, the aging maintenance status (275 days since last release) means you should verify it meets your concurrency and performance needs before committing to a…
Cloudpickle extends Python's standard pickle module to serialize lambda functions, interactively-defined functions and classes, and other constructs that the default pickle cannot handle, making it suitable for cluster computing and remote code execution.
Install it if you need to serialize lambda functions, interactively-defined code, or non-standard Python constructs for cluster computing or distributed execution.
Provides a unified, open()-compatible Python API for streaming large files from remote storage (S3, GCS, Azure, HDFS, SFTP, HTTP) and local filesystems, with transparent compression support.
Install it if you work with large files on cloud storage or remote systems and want to avoid writing boilerplate around multiple SDKs.
Portalocker provides cross-platform file locking with support for exclusive and shared locks, plus Redis-based distributed locks and process-aware PID file locking.
Install it if you need file or process coordination; the optional extras (pywin32, redis) are only required for specific lock types.
Ray is a distributed computing framework that scales Python applications from a single machine to multi-node clusters, providing abstractions for parallel tasks, stateful actors, and shared objects.
See also brickflows · databricks-sql-connector · dataproc-spark-connect · databricks-sdk · nutter · databricks-test · dbx · databricks-cli · databricks-bundles · databricks-api