pbspark
Convert between protobuf messages and pyspark dataframes
Decision gist · record as of 2026-08-14
Yes, if you need protobuf-Spark integration and can tolerate abandoned maintenance. The package is stable with no known vulnerabilities, but expect no updates for future pyspark or protobuf incompatibilities. Suitable for existing projects already committed to this stack; risky for new projects expecting long-term support.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires protobuf message classes generated by protoc; fully qualified module names must match import paths to avoid PicklingError in distributed contexts.
- Low install friction with a pure-Python wheel.
- Maintenance is abandoned—last release was 2023-06-07 with no commits since then—so expect no updates for bugs or dependency conflicts.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package freely with minimal restrictions.
last release 2023-06-07 (1164 days) · last repo commit 2023-06-07 · 23 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 235,040 downloads/mo, #9,012 on PyPI
Alternatives
Verify before relying
from pyspark.sql.session import SparkSession
from pbspark import from_protobuf, to_protobuf
spark = SparkSession.builder.getOrCreate()
df_encoded = spark.createDataFrame([{"value": b"..."}])
df_decoded = df_encoded.select(from_protobuf(df_encoded.value, MessageClass).alias("value"))- Whether abandoned status and lack of recent commits pose compatibility risks with current pyspark or protobuf versions.
- Performance characteristics when handling large message volumes or deeply nested protobuf structures.
- Specific Python version compatibility beyond the stated 3.7–3.11 range.
What it is and what it does
pbspark bridges protobuf and PySpark by providing functions to deserialize protobuf-encoded binary data into Spark StructTypes and re-encode them back. It wraps PySpark UDFs to handle the conversion, with special handling for protobuf's bytes, Timestamp, and int64 types to map them correctly to Spark types. The package offers both column-level operations (from_protobuf, to_protobuf) and DataFrame-level helpers (df_from_protobuf, df_to_protobuf) that can optionally expand struct columns into individual fields.
The MessageConverter class provides a stateful interface for managing conversions and supports custom serializers for non-standard message types. It uses protobuf's MessageToDict internally but overrides its defaults to preserve type fidelity. The package depends on pyspark and protobuf.
Use it for
- Deserialize protobuf-encoded columns in a Spark DataFrame into structured columns for analysis.
- Re-encode expanded DataFrame columns back into protobuf binary format for storage or transmission.
- Define custom serialization logic for specific protobuf message types within a Spark pipeline.
- Convert protobuf data pipelines to Spark for distributed processing without manual schema mapping.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need protobuf-Spark integration and can tolerate abandoned maintenance.
The package is stable with no known vulnerabilities, but expect no updates for future pyspark or protobuf incompatibilities. Suitable for existing projects already committed to this stack; risky for new projects expecting long-term support.
Install
pbspark on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance is abandoned—last release was 2023-06-07 with no commits since then—so expect no updates for bugs or dependency conflicts.
Requires protobuf message classes generated by protoc; fully qualified module names must match import paths to avoid PicklingError in distributed contexts.
License in practice
MIT license is permissive; you can use, modify, and distribute this package freely with minimal restrictions.
Quickstart
from pyspark.sql.session import SparkSession
from pbspark import from_protobuf, to_protobuf
spark = SparkSession.builder.getOrCreate()
df_encoded = spark.createDataFrame([{"value": b"..."}])
df_decoded = df_encoded.select(from_protobuf(df_encoded.value, MessageClass).alias("value"))
Verify before relying
- Whether abandoned status and lack of recent commits pose compatibility risks with current pyspark or protobuf versions.
- Performance characteristics when handling large message volumes or deeply nested protobuf structures.
- Specific Python version compatibility beyond the stated 3.7–3.11 range.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.7,<4.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespysparkprotobuf |
| Maintenance | Abandoned 1,164 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 235,040 / month, #9,012 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaLicense :: OSI Approved :: MIT LicenseProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Database |
Evidence: pbspark-0.9.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “protobuf pyspark conversion”
- pbsparkConverts protobuf messages to and from PySpark DataFrames using UDFs,…
- sparkdanticConverts Pydantic models to PySpark schemas (StructType or JSON…
- pyspark-extensionProvides Python bindings and utilities for Apache Spark, including…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also pyspark-pandas · sparkdantic · sparkaid · protobuf_to_pydantic · bbpb · setuptools-protobuf · pyspark-test · koalas · tinsel