koalas
Koalas: pandas API on Apache Spark
Decision gist · record as of 2026-08-14
No, unless you are locked into Spark 3.1 or below and cannot upgrade. Koalas is dormant (last release October 2021, minimal commits since) and officially deprecated in favor of PySpark's native pandas API layer in Spark 3.2+. For new projects or upgradeable environments, use PySpark directly. For legacy systems on Spark 3.1, Koalas remains functional but will not receive updates.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Apache Spark 3.1 or below; for Spark 3.2+, use PySpark directly.
- Also requires pandas, pyarrow, and numpy as runtime dependencies.
- Low install friction with a pure-Python wheel.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache License 2.0 (permissive), which allows commercial and private use with minimal restrictions, though you must include a copy of the license and state significant changes.
last release 2021-10-19 (1760 days) · last repo commit 2024-03-20 · 3,372 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,084,450 downloads/mo, #4,389 on PyPI
Alternatives
Verify before relying
pip install koalas
import pandas as pd
import pyarrow
import numpy
pdf = pd.DataFrame({'x': range(3), 'y': ['a', 'b', 'b']})
# Convert to Koalas DataFrame and perform operations- Exact import path and API surface for Koalas DataFrame creation and operations
- Whether existing code will continue to work without modification if Spark is upgraded to 3.2+
- Compatibility with Python versions beyond 3.9 (classifiers support up to 3.9)
What it is and what it does
Koalas bridges pandas and Apache Spark by exposing a pandas-compatible DataFrame API that executes on Spark clusters. If you know pandas, you can write similar code against Koalas and have it run on distributed data without learning Spark's native API. The package wraps Spark's distributed execution while mimicking pandas' single-node interface, making it useful for teams that want to scale pandas workflows to big data without rewriting code.
However, Koalas is now in maintenance mode and officially superseded by PySpark's native pandas API layer in Spark 3.2+. The last release was in October 2021, and the repository has seen minimal activity since. For new projects targeting Spark 3.2 or later, the description recommends using PySpark directly instead.
Use it for
- Scale existing pandas code to distributed Spark clusters without rewriting logic or learning Spark's native API
- Write a single codebase that works with pandas on small datasets (for testing) and Koalas on large Spark clusters
- Migrate legacy pandas workflows to big data infrastructure when locked into Spark 3.1 or below
- Prototype data transformations in pandas, then deploy them on Spark using Koalas with minimal code changes
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No, unless you are locked into Spark 3.1 or below and cannot upgrade.
Koalas is dormant (last release October 2021, minimal commits since) and officially deprecated in favor of PySpark's native pandas API layer in Spark 3.2+. For new projects or upgradeable environments, use PySpark directly. For legacy systems on Spark 3.1, Koalas remains functional but will not receive updates.
Install
koalas on PyPI
Before you install
Low install friction with a pure-Python wheel. However, the package is dormant—last release was 2021-10-19 and last commit 2024-03-20—and the description explicitly states it is in maintenance mode, superseded by PySpark's native pandas API layer in Spark 3.2+.
Requires Apache Spark 3.1 or below; for Spark 3.2+, use PySpark directly. Also requires pandas, pyarrow, and numpy as runtime dependencies.
License in practice
Licensed under Apache License 2.0 (permissive), which allows commercial and private use with minimal restrictions, though you must include a copy of the license and state significant changes.
Quickstart
pip install koalas
import pandas as pd
import pyarrow
import numpy
pdf = pd.DataFrame({'x': range(3), 'y': ['a', 'b', 'b']})
# Convert to Koalas DataFrame and perform operations
Verify before relying
- Exact import path and API surface for Koalas DataFrame creation and operations
- Whether existing code will continue to work without modification if Spark is upgraded to 3.2+
- Compatibility with Python versions beyond 3.9 (classifiers support up to 3.9)
Package facts
| License | permissive license permissive |
| Python support | Capped below the current Python release >=3.5,<3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagespandaspyarrownumpy |
| Maintenance | Dormant 1,760 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,084,450 / month, #4,389 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: koalas-1.8.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “pandas api on spark”
- koalasKoalas implements the pandas DataFrame API on top of Apache Spark,…
- pyspark-pandasProvides tools for distributing Pandas DataFrames and Series across…
- pysparkPySpark provides Python bindings to Apache Spark, enabling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also pyspark-pandas · pyspark · pyspark-client · spark-sklearn · graphframes · h2o-pysparkling-3.1 · sagemaker-feature-store-pyspark-3.1 · narwhals · qpd · dataengine