$npx skillfedfor your agent

vortex-data

Python bindings for Vortex, an Apache Arrow-compatible toolkit for working with compressed array data.

With conditionsPyPI Scientific/EngineeringReleased Aug 2026203.4K downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — vortex_data-0.84.0-cp311-abi3-macosx_10_12_x86_64.whl · vortex_data-0.84.0-cp311-abi3-macosx_11_0_arm64.whl · vortex_data-0.84.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
v0.84.0 · released 2026-08-07 · Python >=3.11 · 3 runtime deps: pyarrow, substrait, typing-extensions

Yes, if you work with large columnar datasets and need faster random access or scan performance than Parquet, or if you want to experiment with pluggable compression encodings in an Arrow-compatible format. The package is actively maintained, has no known vulnerabilities, and uses a permissive license. Install friction is moderate due to compiled wheels, but pre-built binaries are available for common platforms. Not necessary if you're already satisfied with Parquet or don't need the performance gains.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.11 or later; compiled wheels available for macOS (x86_64, arm64) and Linux (x86_64, aarch64).
  • Medium install friction due to compiled wheels for multiple platforms (macOS x86_64/arm64, Linux x86_64/aarch64).
  • Active maintenance with releases within the past week and 3138 repository stars.

License · maintenance · safety

permissive license (permissive) — Licensed under Apache License 2.0 (permissive). No restrictions on commercial use, modification, or distribution provided attribution and license terms are included.

last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 3,138 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 203,397 downloads/mo, #9,630 on PyPI

Verify before relying

pip install vortex-data

import vortex
# Use vortex with pyarrow arrays for columnar data operations
  • Specific performance benchmarks (100x faster random access, 10-20x faster scans) are claimed in the description but not independently verified in the fact sheet.
  • File format stability claim (backwards compatible from 0.36.0 onwards) applies to the Rust format; Python API stability guarantees are not explicitly stated.
  • Integration status with Iceberg and other listed systems (Spark, Polars, DuckDB) in the Python bindings specifically.
Same gist for agents: .md · .json

What it is and what it does

Vortex-data is a Python interface to Vortex, a next-generation columnar storage format designed for object-storage-backed data systems. It separates logical schema from physical encoding, allowing pluggable compression strategies (RLE, dictionary, and others) while maintaining zero-copy compatibility with Apache Arrow. The package lets you read, write, and manipulate Vortex files from Python, integrating with the broader Arrow ecosystem.

The format is optimized for random access and scan performance on wide tables with efficient metadata handling. It comes with built-in encodings compatible with Arrow's memory layout and supports cascading compression schemes. The file format itself is considered stable from version 0.36.0 onwards, though the Python library APIs may evolve. Dependencies include pyarrow, substrait, and typing-extensions.

Use it for

  • Store and query large columnar datasets in object storage with faster random access than Parquet.
  • Build data pipelines that need efficient compression without sacrificing read performance.
  • Integrate Vortex files into Arrow-based analytics workflows (DataFusion, DuckDB, Pandas).
  • Experiment with alternative encoding strategies for specialized data types or access patterns.
  • Benchmark columnar formats in performance-critical data systems.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you work with large columnar datasets and need faster random access or scan performance than Parquet, or if you want to experiment with pluggable compression encodings in an Arrow-compatible format.

The package is actively maintained, has no known vulnerabilities, and uses a permissive license. Install friction is moderate due to compiled wheels, but pre-built binaries are available for common platforms. Not necessary if you're already satisfied with Parquet or don't need the performance gains.

Install

vortex-data on PyPI

Before you install

Medium install friction due to compiled wheels for multiple platforms (macOS x86_64/arm64, Linux x86_64/aarch64). Active maintenance with releases within the past week and 3138 repository stars. Requires Python 3.11 or later.

Requires Python 3.11 or later; compiled wheels available for macOS (x86_64, arm64) and Linux (x86_64, aarch64).

License in practice

Licensed under Apache License 2.0 (permissive). No restrictions on commercial use, modification, or distribution provided attribution and license terms are included.

Quickstart

pip install vortex-data

import vortex
# Use vortex with pyarrow arrays for columnar data operations

Verify before relying

  • Specific performance benchmarks (100x faster random access, 10-20x faster scans) are claimed in the description but not independently verified in the fact sheet.
  • File format stability claim (backwards compatible from 0.36.0 onwards) applies to the Rust format; Python API stability guarantees are not explicitly stated.
  • Integration status with Iceberg and other listed systems (Spark, Polars, DuckDB) in the Python bindings specifically.

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.11
Install frictionMedium. Platform-specific wheel
Runtime dependencies
3 packages
pyarrowsubstraittyping-extensions
MaintenanceActively maintained 7 days since the last release
Last repo commit
First released
Downloads203,397 / month, #9,630 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: DatabaseTopic :: File FormatsTopic :: Scientific/Engineering

Evidence: vortex_data-0.84.0-cp311-abi3-macosx_10_12_x86_64.whl; vortex_data-0.84.0-cp311-abi3-macosx_11_0_arm64.whl; vortex_data-0.84.0-cp311-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; vortex_data-0.84.0-cp311-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Tags

Capabilities
columnar file format pythonarrow-compatible data compressionhigh-performance data storagevortex file format bindingscompressed array data toolkit
Topics
columnar-storagearrow-compatiblecompression

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “columnar file format python”

  • vortex-dataVortex-data provides Python bindings to work with Vortex, a columnar…
  • feather-formatFeather-format provides a Python interface to store and load pandas…
  • pyorcPyORC reads and writes Apache ORC files using Python, wrapping the…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also datafusion · pyarrow · pylance · feather-format · total-perspective-vortex · hepconvert · geoarrow-c · arrow-odbc · awkward0 · fastparquet