$npx skillfedfor your agent

sparkaid

Utils for working with Spark

SkipPyPI Build ToolsReleased Aug 2022191.1K downloads / mocopyleft licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — sparkaid-1.0.0-py3-none-any.whl
v1.0.0 · released 2022-08-29 · Python >=3.4 · 1 runtime deps: pyspark

No. The package is abandoned (last release August 2022, no commits since September 2022) and carries a copyleft license that may restrict use in proprietary projects. While it addresses a real problem in Spark schema manipulation, modern Spark versions (3.x+) and the Databricks documentation now provide native or better-maintained alternatives. Install only if you are locked into an older Spark version and cannot upgrade.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires pyspark to be installed and a working Spark environment; package is abandoned and may not be compatible with recent Spark versions.
  • Low install friction with a single runtime dependency on pyspark.
  • However, the package is abandoned—last release was 2022-08-29 and last commit 2022-09-16—so no maintenance or bug fixes should be expected.

License · maintenance · safety

copyleft license (copyleft) — Licensed under LGPLv3+, a copyleft license requiring derivative works to be distributed under the same terms. This may restrict use in proprietary or closed-source projects.

last release 2022-08-29 (1446 days) · last repo commit 2022-09-16 · 7 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 191,085 downloads/mo, #9,892 on PyPI

Verify before relying

pip install sparkaid

from sparkaid import flatten
from pyspark.sql import Row, SparkSession

spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame([Row(structA=Row(field1=10, field2=1.5))])
flattened = flatten(df)
  • Compatibility with Spark versions released after 2022-09-16 is unknown.
  • Whether the snake_case() and json_schema_to_spark_schema() functions mentioned in the changelog are fully documented and stable.
  • Performance characteristics when flattening very large or deeply nested schemas.
Same gist for agents: .md · .json

What it is and what it does

Sparkaid is a utility library for PySpark that simplifies working with DataFrames containing complex nested schemas. It provides functions to flatten StructType columns (removing nesting layers), rename fields within nested structures, and convert JSON schemas to Spark StructType objects. The package addresses common pain points when working with nested data: complex SQL queries, difficulty renaming or casting nested columns, and unnecessary I/O overhead when reading only specific nested columns from Parquet files.

The library's core feature is its flatten() function, which unpacks nested StructType columns into flat columns with configurable separators (e.g., converting {"parent": {"child": "value"}} into {"parent_child": "value"}). Version 1.0.0 introduced a breaking change where flatten() now stops at ArrayType columns by default, requiring explicit configuration to unpack arrays. The package depends solely on pyspark and installs as a pure Python wheel with minimal friction.

Use it for

  • Flatten deeply nested JSON or Parquet data into a single-level schema for simpler SQL queries and easier column access.
  • Rename nested struct fields programmatically without manually reconstructing the entire schema using struct() or cast().
  • Convert JSON schema definitions to Spark StructType objects for schema validation or DataFrame creation.
  • Reduce I/O overhead when reading only specific nested columns from Parquet files by flattening the schema first.
  • Transform snake_case naming conventions across nested column hierarchies in a single operation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No.

The package is abandoned (last release August 2022, no commits since September 2022) and carries a copyleft license that may restrict use in proprietary projects. While it addresses a real problem in Spark schema manipulation, modern Spark versions (3.x+) and the Databricks documentation now provide native or better-maintained alternatives. Install only if you are locked into an older Spark version and cannot upgrade.

Install

sparkaid on PyPI

Before you install

Low install friction with a single runtime dependency on pyspark. However, the package is abandoned—last release was 2022-08-29 and last commit 2022-09-16—so no maintenance or bug fixes should be expected.

Requires pyspark to be installed and a working Spark environment; package is abandoned and may not be compatible with recent Spark versions.

License in practice

Licensed under LGPLv3+, a copyleft license requiring derivative works to be distributed under the same terms. This may restrict use in proprietary or closed-source projects.

Quickstart

pip install sparkaid

from sparkaid import flatten
from pyspark.sql import Row, SparkSession

spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame([Row(structA=Row(field1=10, field2=1.5))])
flattened = flatten(df)

Verify before relying

  • Compatibility with Spark versions released after 2022-09-16 is unknown.
  • Whether the snake_case() and json_schema_to_spark_schema() functions mentioned in the changelog are fully documented and stable.
  • Performance characteristics when flattening very large or deeply nested schemas.

Package facts

Licensecopyleft license copyleft
Python supportSupports the current Python release >=3.4
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
pyspark
MaintenanceAbandoned 1,446 days since the last release
Last repo commit
First released
Downloads191,085 / month, #9,892 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: GNU Lesser General Public License v3 or later (LGPLv3+)Programming Language :: Python :: 3Topic :: Software Development :: Build Tools

Evidence: sparkaid-1.0.0-py3-none-any.whl

Tags

Capabilities
spark dataframe nested schemaflatten spark struct columnsrename nested spark fieldscomplex spark schema utilitiesspark json schema conversionspark dataframe manipulationspark structtype flattening
Topics
spark-dataframeschema-flatteningabandoned
PyPI keywords
SPARKDATAFRAMEMANIPULATE

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “spark dataframe nested schema”

  • sparkaidProvides utilities for working with Spark DataFrames that have…
  • tinselGenerates PySpark DataFrame schemas from Python NamedTuple and…
  • quinnQuinn provides helper methods for PySpark DataFrame validation,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Build Tools packages

packaging Worth it
PyPI · Build Tools · released Aug 2026

Provides reusable utilities for Python packaging interoperability, including version handling, specifiers, markers, requirements, tags, and metadata parsing according to standards like PEP 440 and PEP 425.

Apache-2.0 OR BSD-2-Clausepure Python · 3.9+
2.2Bdownloads / mo
tqdm Worth it
PyPI · Libraries · released Jul 2026

Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.

copyleftpure Python · 3.8+
648.6Mdownloads / mo
pip Worth it
PyPI · Build Tools · released Aug 2026

pip is the standard installer for Python packages, enabling you to download and install packages from the Python Package Index and other indexes into your Python environment.

MITpure Python · 3.10+
617.5Mdownloads / mo
hatchling Worth it
PyPI · Python Modules · released Aug 2026

Hatchling is a standards-compliant Python build backend that handles packaging, metadata, and distribution of Python projects when configured in a project's pyproject.toml file.

MITpure Python · 3.10+
484.2Mdownloads / mo
grpcio-tools Worth it
PyPI · Build Tools · released Jul 2026

Generates Python gRPC service stubs and message classes from Protocol Buffer definitions, enabling developers to build gRPC clients and servers.

Apache-2.0compiled wheel · 3.10+
278.2Mdownloads / mo
pre-commit Worth it
PyPI · Build Tools · released Aug 2026

pre-commit is a framework for installing and running git hooks written in any language before commits are made, automating code quality and validation checks across multi-language projects.

Install it if your team needs consistent, automated validation at commit time.

permissive licensepure Python · 3.10+
179.9Mdownloads / mo

See also quinn · sparkdantic · tinsel · pbspark · chispa · pyspark-test · awkward-pandas · json-tools-rs · dbl-tempo · pyspark-pandas