delta-spark
Python APIs for using Delta Lake with Apache Spark
Decision gist · record as of 2026-08-14
Yes. Delta Lake is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, and installs with low friction. It solves a real problem—ACID reliability in data lakes—and is widely adopted (top 1000 PyPI). Install it if you use Apache Spark and need transaction guarantees or unified batch/streaming semantics; skip it if you don't use Spark or don't need those guarantees.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Apache Spark to be installed and configured; SparkSession must be created with Delta Lake configurations as documented at https://docs.delta.io/latest/delta-intro.html
- Low friction installation as a pure Python wheel.
- Actively maintained with recent releases; repository shows 8939 stars and last commit on 2026-08-13.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing use in commercial and proprietary projects with minimal restrictions beyond attribution and liability disclaimers.
last release 2026-07-08 (37 days) · last repo commit 2026-08-13 · 8,939 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 38,466,157 downloads/mo, #715 on PyPI
Alternatives
Verify before relying
pip install delta-spark
from delta.tables import DeltaTable
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("delta-example").getOrCreate()
delta_table = DeltaTable.forPath(spark, "/path/to/delta/table")- Specific Delta Lake version compatibility with different Spark versions (not stated in fact sheet)
- Whether importlib_metadata is used for runtime feature detection or is a build-time dependency only
What it is and what it does
Delta Lake is a storage layer that adds ACID transaction guarantees, reliable metadata handling, and unified batch/streaming semantics to Apache Spark data lakes. The delta-spark package provides the Python API bindings to interact with Delta Lake tables from PySpark code, allowing you to read, write, and manage Delta tables with transactional guarantees and schema enforcement.
It integrates directly with Spark's DataFrame API and runs on top of existing data lake storage (HDFS, S3, etc.), making it a drop-in reliability layer for Spark workloads. The package depends on pyspark and importlib_metadata, and requires Python 3.10 or later. Configuration is handled through SparkSession setup rather than the package itself.
Use it for
- Build reliable ETL pipelines with ACID guarantees for data consistency across batch and streaming jobs
- Manage evolving data schemas with Delta Lake's schema enforcement and evolution capabilities
- Implement time-travel queries to audit data changes or recover from accidental overwrites
- Unify batch and streaming workloads on the same table without data consistency issues
- Replace data lake governance gaps with transaction logs and metadata versioning
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Delta Lake is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, and installs with low friction. It solves a real problem—ACID reliability in data lakes—and is widely adopted (top 1000 PyPI). Install it if you use Apache Spark and need transaction guarantees or unified batch/streaming semantics; skip it if you don't use Spark or don't need those guarantees.
Install
delta-spark on PyPI
Before you install
Low friction installation as a pure Python wheel. Actively maintained with recent releases; repository shows 8939 stars and last commit on 2026-08-13. Requires Python 3.10 or later and pyspark as a runtime dependency.
Requires Apache Spark to be installed and configured; SparkSession must be created with Delta Lake configurations as documented at https://docs.delta.io/latest/delta-intro.html
License in practice
Licensed under Apache-2.0 (permissive), allowing use in commercial and proprietary projects with minimal restrictions beyond attribution and liability disclaimers.
Quickstart
pip install delta-spark
from delta.tables import DeltaTable
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("delta-example").getOrCreate()
delta_table = DeltaTable.forPath(spark, "/path/to/delta/table")
Verify before relying
- Specific Delta Lake version compatibility with different Spark versions (not stated in fact sheet)
- Whether importlib_metadata is used for runtime feature detection or is a build-time dependency only
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespysparkimportlib_metadata |
| Maintenance | Actively maintained 37 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 38,466,157 / month, #715 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Software Development :: Libraries :: Python ModulesTyping :: Typed |
Evidence: delta_spark-4.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “delta lake python spark”
- delta-sparkProvides Python APIs for Delta Lake, an ACID-transactional storage…
- hops-deltalakeReads and writes Delta Lake tables with native Python bindings backed…
- delta-sharingDelta Sharing is a Python client library for the Delta Sharing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Python Modules packages
Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.
Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.
Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.
PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.
Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…
Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.
See also delta-sharing · deltalite · deltalake · hops-deltalake · dbt-databricks · pyspark-pandas · lakefs · delta-kernel-rust-sharing-wrapper · pyspark · dbl-tempo