$npx skillfedfor your agent

koalas

Koalas: pandas API on Apache Spark

SkipPyPI Information AnalysisReleased Oct 20211.1M downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — koalas-1.8.2-py3-none-any.whl
v1.8.2 · released 2021-10-19 · Python >=3.5,<3.10 · 3 runtime deps: pandas, pyarrow, numpy

No, unless you are locked into Spark 3.1 or below and cannot upgrade. Koalas is dormant (last release October 2021, minimal commits since) and officially deprecated in favor of PySpark's native pandas API layer in Spark 3.2+. For new projects or upgradeable environments, use PySpark directly. For legacy systems on Spark 3.1, Koalas remains functional but will not receive updates.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Apache Spark 3.1 or below; for Spark 3.2+, use PySpark directly.
  • Also requires pandas, pyarrow, and numpy as runtime dependencies.
  • Low install friction with a pure-Python wheel.

License · maintenance · safety

permissive license (permissive) — Licensed under Apache License 2.0 (permissive), which allows commercial and private use with minimal restrictions, though you must include a copy of the license and state significant changes.

last release 2021-10-19 (1760 days) · last repo commit 2024-03-20 · 3,372 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,084,450 downloads/mo, #4,389 on PyPI

Verify before relying

pip install koalas

import pandas as pd
import pyarrow
import numpy

pdf = pd.DataFrame({'x': range(3), 'y': ['a', 'b', 'b']})
# Convert to Koalas DataFrame and perform operations
  • Exact import path and API surface for Koalas DataFrame creation and operations
  • Whether existing code will continue to work without modification if Spark is upgraded to 3.2+
  • Compatibility with Python versions beyond 3.9 (classifiers support up to 3.9)
Same gist for agents: .md · .json

What it is and what it does

Koalas bridges pandas and Apache Spark by exposing a pandas-compatible DataFrame API that executes on Spark clusters. If you know pandas, you can write similar code against Koalas and have it run on distributed data without learning Spark's native API. The package wraps Spark's distributed execution while mimicking pandas' single-node interface, making it useful for teams that want to scale pandas workflows to big data without rewriting code.

However, Koalas is now in maintenance mode and officially superseded by PySpark's native pandas API layer in Spark 3.2+. The last release was in October 2021, and the repository has seen minimal activity since. For new projects targeting Spark 3.2 or later, the description recommends using PySpark directly instead.

Use it for

  • Scale existing pandas code to distributed Spark clusters without rewriting logic or learning Spark's native API
  • Write a single codebase that works with pandas on small datasets (for testing) and Koalas on large Spark clusters
  • Migrate legacy pandas workflows to big data infrastructure when locked into Spark 3.1 or below
  • Prototype data transformations in pandas, then deploy them on Spark using Koalas with minimal code changes

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No, unless you are locked into Spark 3.1 or below and cannot upgrade.

Koalas is dormant (last release October 2021, minimal commits since) and officially deprecated in favor of PySpark's native pandas API layer in Spark 3.2+. For new projects or upgradeable environments, use PySpark directly. For legacy systems on Spark 3.1, Koalas remains functional but will not receive updates.

Install

koalas on PyPI

Before you install

Low install friction with a pure-Python wheel. However, the package is dormant—last release was 2021-10-19 and last commit 2024-03-20—and the description explicitly states it is in maintenance mode, superseded by PySpark's native pandas API layer in Spark 3.2+.

Requires Apache Spark 3.1 or below; for Spark 3.2+, use PySpark directly. Also requires pandas, pyarrow, and numpy as runtime dependencies.

License in practice

Licensed under Apache License 2.0 (permissive), which allows commercial and private use with minimal restrictions, though you must include a copy of the license and state significant changes.

Quickstart

pip install koalas

import pandas as pd
import pyarrow
import numpy

pdf = pd.DataFrame({'x': range(3), 'y': ['a', 'b', 'b']})
# Convert to Koalas DataFrame and perform operations

Verify before relying

  • Exact import path and API surface for Koalas DataFrame creation and operations
  • Whether existing code will continue to work without modification if Spark is upgraded to 3.2+
  • Compatibility with Python versions beyond 3.9 (classifiers support up to 3.9)

Package facts

Licensepermissive license permissive
Python supportCapped below the current Python release >=3.5,<3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
pandaspyarrownumpy
MaintenanceDormant 1,760 days since the last release
Last repo commit
First released
Downloads1,084,450 / month, #4,389 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: koalas-1.8.2-py3-none-any.whl

Tags

Capabilities
pandas api on sparkdistributed dataframe sparkspark pandas compatibilitybig data dataframe pythonspark dataframe wrapperpandas to spark migrationdistributed pandas alternative
Topics
spark-integrationdataframe-apideprecated

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pandas api on spark”

  • koalasKoalas implements the pandas DataFrame API on top of Apache Spark,…
  • pyspark-pandasProvides tools for distributing Pandas DataFrames and Series across…
  • pysparkPySpark provides Python bindings to Apache Spark, enabling…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also pyspark-pandas · pyspark · pyspark-client · spark-sklearn · graphframes · h2o-pysparkling-3.1 · sagemaker-feature-store-pyspark-3.1 · narwhals · qpd · dataengine