graphframes
GraphFrames: DataFrame-based Graphs
Decision gist · record as of 2026-08-14
Yes, if you have Apache Spark and need to run graph algorithms on data too large for a single machine. The library is actively maintained, has low install friction, and carries a permissive MIT license. However, verify that your Spark version and Python environment are compatible, since the latest release date and classifier information suggest the package may not have been updated for very recent versions.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Apache Spark to be installed and configured; GraphFrames is a Spark library, not a standalone graph tool.
- Low install friction with a pure-wheel distribution.
- The package is actively maintained with a recent commit on 2026-08-12, though the latest release date shown is 2018-12-05.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most production environments.
last release 2018-12-05 (2809 days) · last repo commit 2026-08-12 · 1,201 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,683,394 downloads/mo, #2,940 on PyPI
Alternatives
Verify before relying
pip install graphframes
from graphframes import GraphFrame
g = GraphFrame(nodes_df, edges_df)
print(g.inDegrees.show())- Whether the package supports current Python versions beyond 3.6, given classifiers list only up to 3.6
- Current Apache Spark version compatibility, as the last release is dated 2018-12-05
- Whether numpy and nose are truly runtime dependencies or only test/build-time requirements
What it is and what it does
GraphFrames is a Python library that brings graph processing to Apache Spark's distributed computing framework. It sits on top of Spark's DataFrame API, letting you represent graphs as vertex and edge DataFrames, then run graph algorithms—like PageRank, connected components, shortest paths, and motif finding—across a cluster. The package combines relational queries with graph traversals, so you can filter and join graph data using SQL-like syntax while leveraging Spark's optimizer for performance.
Typical use involves creating vertex and edge DataFrames, constructing a GraphFrame object, then calling built-in algorithms or writing custom logic with Pregel and message-passing APIs. It's designed for scenarios where your graph is too large for a single machine and you need both graph-specific operations and the flexibility to combine them with relational transformations.
Use it for
- Entity resolution at scale by connecting similar records and running connected components to group duplicates
- Fraud detection in large transaction networks using cycle detection and K-Core algorithm
- Social network analysis such as ranking search results with distributed PageRank or finding independent sets for marketing campaigns
- Compliance analytics using shortest-path algorithms and motif analysis to detect suspicious patterns
- Knowledge graph construction and querying with property graph models and relational joins
- Graph clustering and community detection on massive networks using label propagation
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have Apache Spark and need to run graph algorithms on data too large for a single machine.
The library is actively maintained, has low install friction, and carries a permissive MIT license. However, verify that your Spark version and Python environment are compatible, since the latest release date and classifier information suggest the package may not have been updated for very recent versions.
Install
graphframes on PyPI
Before you install
Low install friction with a pure-wheel distribution. The package is actively maintained with a recent commit on 2026-08-12, though the latest release date shown is 2018-12-05.
Requires Apache Spark to be installed and configured; GraphFrames is a Spark library, not a standalone graph tool.
License in practice
MIT license permits commercial and private use with minimal restrictions, making it suitable for most production environments.
Quickstart
pip install graphframes
from graphframes import GraphFrame
g = GraphFrame(nodes_df, edges_df)
print(g.inDegrees.show())
Verify before relying
- Whether the package supports current Python versions beyond 3.6, given classifiers list only up to 3.6
- Current Apache Spark version compatibility, as the last release is dated 2018-12-05
- Whether numpy and nose are truly runtime dependencies or only test/build-time requirements
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesnumpynose |
| Maintenance | Actively maintained 2,809 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,683,394 / month, #2,940 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.5Programming Language :: Python :: 3.6Topic :: Scientific/Engineering |
Evidence: graphframes-0.6-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “distributed graph algorithms spark”
- graphframesGraphFrames provides distributed graph processing and analytics on…
- graphframes-pyGraphFrames Python wrapper provides graph processing and analysis on…
- pysparkPySpark provides Python bindings to Apache Spark, enabling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also graphframes-py · pyspark · pyspark-pandas · scikit-network · pyspark-client · grandalf · koalas · graspologic · nebula3-python · networkx