h3-pyspark
PySpark bindings for H3, a hierarchical hexagonal geospatial indexing system
Decision gist · record as of 2026-08-14
Yes, if you are already using PySpark and need distributed H3 indexing—it has no runtime dependencies, low install friction, and a permissive license. However, be aware that the package is dormant; verify compatibility with your PySpark and H3 versions before committing to production. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires PySpark to be installed and a running Spark environment; requires Python >=3.6.
- Low friction installation as a pure-Python wheel.
- Maintenance is dormant—last release was 2022-03-10, with no updates since, though the repository remains active and unarchived.
License · maintenance · safety
permissive license (permissive) — Licensed under MIT (permissive), allowing use in proprietary and open-source projects without significant restrictions.
last release 2022-03-10 (1618 days) · last repo commit 2024-03-26 · 33 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 225,618 downloads/mo, #9,218 on PyPI
Alternatives
Verify before relying
pip install h3-pyspark
from pyspark.sql import SparkSession, functions as F
import h3_pyspark
spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame([{"lat": 37.769377, "lng": -122.388903, "resolution": 9}])
df = df.withColumn('h3_9', h3_pyspark.geo_to_h3('lat', 'lng', 'resolution'))
df.show()- Whether the package works with recent PySpark versions (last release predates significant Spark API changes).
- Compatibility with modern H3 core library versions and whether h3-py dependency is pinned or flexible.
- Performance characteristics on large-scale distributed workloads compared to native Spark geospatial functions.
What it is and what it does
h3-pyspark wraps Uber's H3 hierarchical hexagonal indexing system for use in PySpark DataFrames, enabling you to convert latitude/longitude coordinates into H3 cell identifiers and index complex geometries (points, polygons, multipolygons) as sets of H3 cells at a chosen resolution. It extends the vanilla H3 library with PySpark-native operations for spatial indexing, k-ring buffering, and spatial joins—allowing you to bucket and cluster geometries efficiently across a distributed cluster.
The package assumes GeoJSON representation of geometries and H3 cells as string columns, making it a natural fit for pipelines that already work with GeoJSON. It is most useful for approximate spatial joins and distance-based bucketing, though results are candidates rather than exact matches and should be validated with a secondary distance check if precision is required.
Use it for
- Index geographic features (buildings, roads, regions) into H3 cells for distributed spatial bucketing and clustering.
- Perform approximate spatial joins between two large datasets by indexing both on H3 and joining on cell identity.
- Generate buffered spatial indexes around geometries using k-ring operations for distance-based queries.
- Organize geospatial data into hierarchical hexagonal grids for efficient map visualization and aggregation.
- Implement distance joins by combining H3 indexing with secondary distance validation using a UDF.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are already using PySpark and need distributed H3 indexing—it has no runtime dependencies, low install friction, and a permissive license.
However, be aware that the package is dormant; verify compatibility with your PySpark and H3 versions before committing to production. No known security vulnerabilities.
Install
h3-pyspark on PyPI
Before you install
Low friction installation as a pure-Python wheel. Maintenance is dormant—last release was 2022-03-10, with no updates since, though the repository remains active and unarchived. No runtime dependencies to manage.
Requires PySpark to be installed and a running Spark environment; requires Python >=3.6.
License in practice
Licensed under MIT (permissive), allowing use in proprietary and open-source projects without significant restrictions.
Quickstart
pip install h3-pyspark
from pyspark.sql import SparkSession, functions as F
import h3_pyspark
spark = SparkSession.builder.getOrCreate()
df = spark.createDataFrame([{"lat": 37.769377, "lng": -122.388903, "resolution": 9}])
df = df.withColumn('h3_9', h3_pyspark.geo_to_h3('lat', 'lng', 'resolution'))
df.show()
Verify before relying
- Whether the package works with recent PySpark versions (last release predates significant Spark API changes).
- Compatibility with modern H3 core library versions and whether h3-py dependency is pinned or flexible.
- Performance characteristics on large-scale distributed workloads compared to native Spark geospatial functions.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.6 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Dormant 1,618 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 225,618 / month, #9,218 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: h3_pyspark-1.2.6-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pyspark geospatial indexing”
- h3-pysparkProvides PySpark bindings for H3, enabling hexagonal geospatial…
- h3h3 provides Python bindings to Uber's H3 geospatial indexing library,…
- rasterixRasterix provides tools for analyzing raster data within Xarray…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also h3 · s2sphere · pydeck · geopandas · dask-geopandas · pyspark-extension · pyspark-hnsw · pygeoif · topojson · cmeel-octomap