pyarrow
Python library for Apache Arrow
Decision gist · record as of 2026-08-14
Yes. pyarrow is actively maintained, widely used (top 100 PyPI), has no known vulnerabilities, and provides essential functionality for modern data engineering. Install friction is manageable with pre-built wheels for Python 3.10–3.14. Recommended for any workflow involving Arrow format data or requiring efficient columnar interchange.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- On Windows, may require Visual C++ Redistributable for Visual Studio if wheel import fails.
- Medium install friction due to compiled C++ dependencies, but pre-built wheels cover Python 3.10–3.14 and common platforms (macOS, Linux, Windows).
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; attribution required but no copyleft obligations.
last release 2026-08-10 (4 days) · last repo commit 2026-08-14 · 17,021 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 432,924,469 downloads/mo, #95 on PyPI
Alternatives
Verify before relying
pip install pyarrow
import pyarrow as pa
table = pa.table({'col1': [1, 2, 3], 'col2': ['a', 'b', 'c']})
print(table)- Specific performance characteristics for particular workloads
- Memory overhead of columnar format vs. row-oriented alternatives
- Supported Arrow IPC and file format versions
What it is and what it does
pyarrow is the Python interface to Apache Arrow, a cross-language columnar in-memory data format designed for efficient analytics. It exposes Arrow's C++ implementation to Python, enabling fast data serialization, deserialization, and data sharing with libraries like pandas and NumPy. The package handles Arrow's native data types and provides tools for reading and writing Arrow files and streams.
Developers use pyarrow to build data pipelines that need high-performance columnar storage, language interoperability, or efficient data interchange between Python and other systems. It's commonly embedded in larger data tools but can also be used directly for Arrow format I/O and data transformation tasks.
Use it for
- Reading and writing Apache Arrow IPC and file formats from Python
- Data sharing between Python libraries for memory-efficient workflows
- Building polyglot data pipelines that exchange columnar data with non-Python services
- High-performance columnar analytics on large datasets
- Serializing Python data structures to Arrow format for distributed computing
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
pyarrow is actively maintained, widely used (top 100 PyPI), has no known vulnerabilities, and provides essential functionality for modern data engineering. Install friction is manageable with pre-built wheels for Python 3.10–3.14. Recommended for any workflow involving Arrow format data or requiring efficient columnar interchange.
Install
pyarrow on PyPI
Before you install
Medium install friction due to compiled C++ dependencies, but pre-built wheels cover Python 3.10–3.14 and common platforms (macOS, Linux, Windows). Active maintenance with release 4 days old and 17021 repository stars.
Requires Python 3.10 or later. On Windows, may require Visual C++ Redistributable for Visual Studio if wheel import fails.
License in practice
Apache-2.0 permissive license allows commercial and private use with minimal restrictions; attribution required but no copyleft obligations.
Quickstart
pip install pyarrow
import pyarrow as pa
table = pa.table({'col1': [1, 2, 3], 'col2': ['a', 'b', 'c']})
print(table)
Verify before relying
- Specific performance characteristics for particular workloads
- Memory overhead of columnar format vs. row-oriented alternatives
- Supported Arrow IPC and file format versions
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 4 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 432,924,469 / month, #95 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: Free Threading :: 2 - Beta |
Evidence: pyarrow-25.0.1-cp310-cp310-macosx_12_0_arm64.whl; pyarrow-25.0.1-cp310-cp310-macosx_12_0_x86_64.whl; pyarrow-25.0.1-cp310-cp310-manylinux_2_28_aarch64.whl; pyarrow-25.0.1-cp310-cp310-manylinux_2_28_x86_64.whl; pyarrow-25.0.1-cp310-cp310-musllinux_1_2_aarch64.whl; pyarrow-25.0.1-cp310-cp310-musllinux_1_2_x86_64.whl; pyarrow-25.0.1-cp310-cp310-win_amd64.whl; pyarrow-25.0.1-cp311-cp311-macosx_12_0_arm64.whl; pyarrow-25.0.1-cp311-cp311-macosx_12_0_x86_64.whl; pyarrow-25.0.1-cp311-cp311-manylinux_2_28_aarch64.whl; pyarrow-25.0.1-cp311-cp311-manylinux_2_28_x86_64.whl; pyarrow-25.0.1-cp311-cp311-musllinux_1_2_aarch64.whl; pyarrow-25.0.1-cp311-cp311-musllinux_1_2_x86_64.whl; pyarrow-25.0.1-cp311-cp311-win_amd64.whl; pyarrow-25.0.1-cp312-cp312-macosx_12_0_arm64.whl; pyarrow-25.0.1-cp312-cp312-macosx_12_0_x86_64.whl; pyarrow-25.0.1-cp312-cp312-manylinux_2_28_aarch64.whl; pyarrow-25.0.1-cp312-cp312-manylinux_2_28_x86_64.whl; pyarrow-25.0.1-cp312-cp312-musllinux_1_2_aarch64.whl; pyarrow-25.0.1-cp312-cp312-musllinux_1_2_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “columnar data format python”
- pyarrowpyarrow provides Python bindings to Apache Arrow's C++ libraries for…
- pylancePylance is a Python wrapper for the Lance columnar data format,…
- vortex-dataVortex-data provides Python bindings to work with Vortex, a columnar…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
Snowflake SQLAlchemy is a SQLAlchemy dialect that enables SQLAlchemy applications to connect to and query Snowflake databases using standard SQLAlchemy ORM and Core APIs.
Install it if you are building a Python application that needs to connect to Snowflake and prefer SQLAlchemy's abstraction layer over raw SQL or the connector API.
See also datafusion · feather-format · pydantic-to-pyarrow · pylance · vortex-data · pyarrow-hotfix · arrow-odbc · geoarrow-pyarrow · adbc-driver-sqlite · geoarrow-pandas