pydantic-spark
Converting pydantic classes to spark schemas
Decision gist · record as of 2026-08-14
Yes, if you actively use both Pydantic and Spark and want to avoid duplicating schema definitions. The low install friction and permissive license make it a low-risk addition. However, be aware that the project is dormant—no updates since late 2023—so you should verify compatibility with your current Pydantic and Spark versions before relying on it in production, and plan to maintain a fork if critical bugs emerge.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires pydantic as a runtime dependency; Spark itself is not listed as a dependency, so you must have it available in your environment separately.
- Low install friction with a single runtime dependency on pydantic.
- Maintenance is dormant—last commit was 2024-03-04 and no releases since 2023-11-24—so expect no active bug fixes or feature development, though the codebase remains archived and available.
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing commercial and private use with minimal restrictions; you may use this package freely in most contexts.
last release 2023-11-24 (994 days) · last repo commit 2024-03-04 · 26 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 199,617 downloads/mo, #9,702 on PyPI
Alternatives
Verify before relying
from pydantic_spark.base import SparkBase
class TestModel(SparkBase):
key1: str
key2: int
schema_dict = TestModel.spark_schema()- Whether Spark is an implicit peer dependency or truly optional at runtime.
- Compatibility with recent Pydantic v2 major version changes and their breaking API shifts.
- Whether the coerce_type feature and other advanced options are production-ready or experimental.
What it is and what it does
pydantic-spark is a lightweight bridge between Pydantic's Python type system and Apache Spark's schema representation. It lets you define data models as Pydantic classes and automatically generate the corresponding Spark schema dictionaries, or reverse the process by generating Python code from an existing Spark schema. The library extends Pydantic's BaseModel through a SparkBase class and adds a spark_schema() method that outputs a schema-compatible dictionary.
The package is designed for data engineers and Python developers working with Spark who want to avoid manually writing schema definitions or keeping Pydantic models and Spark schemas in sync. It includes a coerce_type option for field-level type conversion during schema generation. With only pydantic as a runtime dependency, installation is straightforward, though the project is currently dormant—last updated in late 2023—so new features or maintenance are unlikely.
Use it for
- Define Spark DataFrame schemas using Pydantic classes, ensuring type safety and validation in Python before writing to Spark.
- Generate boilerplate Spark schema code from existing Pydantic models to reduce manual schema definition work.
- Reverse-engineer Python Pydantic classes from Spark schemas to keep data contracts synchronized across systems.
- Apply field-level type coercion rules during schema generation when Pydantic types need to map to different Spark types.
- Validate and document data pipelines by using Pydantic's validation alongside Spark's schema enforcement.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you actively use both Pydantic and Spark and want to avoid duplicating schema definitions.
The low install friction and permissive license make it a low-risk addition. However, be aware that the project is dormant—no updates since late 2023—so you should verify compatibility with your current Pydantic and Spark versions before relying on it in production, and plan to maintain a fork if critical bugs emerge.
Install
pydantic-spark on PyPI
Before you install
Low install friction with a single runtime dependency on pydantic. Maintenance is dormant—last commit was 2024-03-04 and no releases since 2023-11-24—so expect no active bug fixes or feature development, though the codebase remains archived and available.
Requires pydantic as a runtime dependency; Spark itself is not listed as a dependency, so you must have it available in your environment separately.
License in practice
MIT license is permissive, allowing commercial and private use with minimal restrictions; you may use this package freely in most contexts.
Quickstart
from pydantic_spark.base import SparkBase
class TestModel(SparkBase):
key1: str
key2: int
schema_dict = TestModel.spark_schema()
Verify before relying
- Whether Spark is an implicit peer dependency or truly optional at runtime.
- Compatibility with recent Pydantic v2 major version changes and their breaking API shifts.
- Whether the coerce_type feature and other advanced options are production-ready or experimental.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.8,<4.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagepydantic |
| Maintenance | Dormant 994 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 199,617 / month, #9,702 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: pydantic_spark-1.0.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pydantic to spark schema”
- pydantic-sparkConverts Pydantic class definitions to Apache Spark schemas and…
- sparkdanticConverts Pydantic models to PySpark schemas (StructType or JSON…
- sparkaidProvides utilities for working with Spark DataFrames that have…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also pydantic-avro · sparkdantic · jsonschema-pydantic · dydantic · jsonschema-pydantic-converter · py-avro-schema · dataclasses-avroschema · django-pydantic-field · schema · pydantic-to-pyarrow