$npx skillfedfor your agent

petastorm

Petastorm is a library enabling the use of Parquet storage from Tensorflow, Pytorch, and other Python-based ML training frameworks.

With conditionsPyPI Artificial IntelligenceReleased Jan 2026199.6K downloads / moApache License, Version 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — petastorm-0.13.1-py2.py3-none-any.whl
v0.13.1 · released 2026-01-02 · Python >=3 · 13 runtime deps: dill, diskcache, future, numpy, packaging, pandas, psutil, pyspark

Yes, if you are already using Parquet for data storage and training with TensorFlow, PyTorch, or PySpark. The low install friction and permissive license make it a straightforward addition. However, note the aging maintenance status (224 days since last release)—verify compatibility with your specific framework versions before committing to production use. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a Parquet dataset already created in Petastorm format; generating one requires PySpark and the Unischema API shown in the documentation.
  • Low friction install with a pure-wheel distribution and 13 runtime dependencies already packaged.
  • Maintenance is aging—last release was 224 days ago—but the repository remains active and unarchived with 1890 stars.

License · maintenance · safety

Apache License, Version 2.0 (permissive) — Apache License 2.0 is permissive, allowing commercial and private use with minimal restrictions; you must retain license notices in distributions.

last release 2026-01-02 (224 days) · last repo commit 2026-01-02 · 1,890 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 199,587 downloads/mo, #9,705 on PyPI

Verify before relying

pip install petastorm

from petastorm import make_reader

with make_reader('file:///path/to/dataset') as reader:
    for row in reader:
        print(row)
  • Whether the aging maintenance status (224 days since last release) affects compatibility with recent TensorFlow or PyTorch versions.
  • Performance characteristics and scalability limits for very large distributed datasets.
  • Current state of optional dependencies (tf, tf_gpu, torch, opencv) and their version compatibility.
Same gist for agents: .md · .json

What it is and what it does

Petastorm is a data access library that bridges Apache Parquet storage and popular Python machine learning frameworks. It was developed at Uber ATG to enable efficient, distributed training of deep learning models directly from Parquet datasets without requiring intermediate format conversions. The library handles the schema mapping between Parquet and framework-specific types (TensorFlow, PyTorch, PySpark), and provides a unified reader interface that supports selective column access, row filtering, shuffling, and partitioning for multi-GPU training.

You use Petastorm in two phases: first, generate a Parquet dataset using PySpark with the Unischema API to define field types, shapes, and compression codecs; then, read from that dataset using the Reader class or framework-specific adapters (tf_tensors for TensorFlow, DataLoader for PyTorch). The library handles parallelism internally via threads, processes, or single-threaded modes, and supports local caching to reduce repeated I/O.

Use it for

  • Train TensorFlow or PyTorch models on large Parquet datasets stored on HDFS or local filesystems without loading entire datasets into memory.
  • Generate Parquet datasets from raw data using PySpark, then iterate over them in Python ML training loops with automatic schema validation.
  • Distribute training across multiple GPUs by partitioning Petastorm datasets and reading different partitions on different workers.
  • Apply row-level filtering and selective column reads to reduce I/O when training on subsets of large datasets.
  • Compress image and array data using standard codecs (JPEG, PNG) or custom codecs within a single Parquet-backed dataset.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are already using Parquet for data storage and training with TensorFlow, PyTorch, or PySpark.

The low install friction and permissive license make it a straightforward addition. However, note the aging maintenance status (224 days since last release)—verify compatibility with your specific framework versions before committing to production use. No known security vulnerabilities.

Install

petastorm on PyPI

Before you install

Low friction install with a pure-wheel distribution and 13 runtime dependencies already packaged. Maintenance is aging—last release was 224 days ago—but the repository remains active and unarchived with 1890 stars.

Requires a Parquet dataset already created in Petastorm format; generating one requires PySpark and the Unischema API shown in the documentation.

License in practice

Apache License 2.0 is permissive, allowing commercial and private use with minimal restrictions; you must retain license notices in distributions.

Quickstart

pip install petastorm

from petastorm import make_reader

with make_reader('file:///path/to/dataset') as reader:
    for row in reader:
        print(row)

Verify before relying

  • Whether the aging maintenance status (224 days since last release) affects compatibility with recent TensorFlow or PyTorch versions.
  • Performance characteristics and scalability limits for very large distributed datasets.
  • Current state of optional dependencies (tf, tf_gpu, torch, opencv) and their version compatibility.

Package facts

LicenseApache License, Version 2.0 permissive
Python supportSupports the current Python release >=3
Install frictionLow. Pure-Python wheel
Runtime dependencies
13 packages
dilldiskcachefuturenumpypackagingpandaspsutilpysparkpyzmqpyarrowsixfsspecsetuptools
MaintenanceAging 224 days since the last release
Last repo commit
First released
Downloads199,587 / month, #9,705 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Environment :: ConsoleEnvironment :: Web EnvironmentIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3.4Programming Language :: Python :: 3.5Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8

Evidence: petastorm-0.13.1-py2.py3-none-any.whl

Tags

Capabilities
parquet data loader for deep learningtensorflow pytorch data pipelinedistributed machine learning dataset accessparquet to tensorflow pytorchspark dataset for ml trainingefficient data loading ml frameworksparquet reader machine learning
Topics
data-loadingparquetdistributed-training

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “parquet data loader for deep learning”

  • petastormPetastorm enables direct training of deep learning models from Apache…
  • mltableMLTable provides fast, flexible data loading and schema abstraction…
  • torchgeoTorchGeo provides PyTorch datasets, samplers, transforms, and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also litdata · datasets · webdataset · spark-sklearn · mosaicml-streaming · torch · tensorflow-datasets · pyspark-huggingface · raydp · keras-nightly