parquet
Python support for Parquet file format
Decision gist · record as of 2026-08-14
No. The package is abandoned (last commit 2021-10-26, no release since 2020-04-30) and was only tested on Python 2.7, 3.6, and 3.7—all now end-of-life. Compatibility with modern Python versions is unverified, and critical Parquet features remain unimplemented. For production use, prefer actively maintained alternatives like pyarrow or fastparquet.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires thriftpy2 and backports.csv as runtime dependencies; compatibility with Python versions newer than 3.7 is unverified.
- Low install friction with only two runtime dependencies.
- However, the package is abandoned—last commit was 2021-10-26 and no release since 2020-04-30.
License · maintenance · safety
Apache License 2.0 (permissive) — Licensed under Apache License 2.0 (permissive), so you can use it freely in commercial and open-source projects without significant legal constraints.
last release 2020-04-30 (2297 days) · last repo commit 2021-10-26 · 362 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 293,204 downloads/mo, #7,956 on PyPI
Alternatives
Verify before relying
import parquet
with open("test.parquet") as fo:
for row in parquet.DictReader(fo, columns=['foo', 'bar']):
print(row)- Whether the package works reliably on Python 3.8 and later versions.
- Current state of nested data support and which parquet-format features remain unimplemented.
- Performance characteristics compared to modern alternatives.
- Whether snappy compression support (optional extra) is still functional.
What it is and what it does
parquet is a pure-Python implementation of the Apache Parquet file format with read-only support. It provides both a command-line tool (`parquet` command) for inspecting Parquet files and a programmatic API with DictReader and reader classes similar to Python's csv module. The package is useful for debugging and quick viewing of Parquet data without the overhead of a JVM, and it can read data files from the parquet-compatibility project.
The package depends on thriftpy2 for Thrift serialization and backports.csv for CSV output. However, it has not been actively maintained since 2020, and its test coverage was limited to Python 2.7, 3.6, and 3.7. Not all Parquet format features have been implemented—notably nested data handling and deprecated bitpacking remain incomplete. Performance has not been optimized.
Use it for
- Quickly inspect Parquet file contents from the command line without installing Java or Spark.
- Convert Parquet data to JSON or TSV format for integration with non-Parquet tools.
- Read Parquet files programmatically in Python for data debugging and exploratory analysis.
- Extract specific columns from Parquet files using the DictReader or reader interface.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
The package is abandoned (last commit 2021-10-26, no release since 2020-04-30) and was only tested on Python 2.7, 3.6, and 3.7—all now end-of-life. Compatibility with modern Python versions is unverified, and critical Parquet features remain unimplemented. For production use, prefer actively maintained alternatives like pyarrow or fastparquet.
Install
parquet on PyPI
Before you install
Low install friction with only two runtime dependencies. However, the package is abandoned—last commit was 2021-10-26 and no release since 2020-04-30. It was tested on Python 2.7, 3.6, and 3.7, which are now end-of-life versions; compatibility with modern Python is unverified.
Requires thriftpy2 and backports.csv as runtime dependencies; compatibility with Python versions newer than 3.7 is unverified.
License in practice
Licensed under Apache License 2.0 (permissive), so you can use it freely in commercial and open-source projects without significant legal constraints.
Quickstart
import parquet
with open("test.parquet") as fo:
for row in parquet.DictReader(fo, columns=['foo', 'bar']):
print(row)
Verify before relying
- Whether the package works reliably on Python 3.8 and later versions.
- Current state of nested data support and which parquet-format features remain unimplemented.
- Performance characteristics compared to modern alternatives.
- Whether snappy compression support (optional extra) is still functional.
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesthriftpy2backports.csv |
| Maintenance | Abandoned 2,297 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 293,204 / month, #7,956 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersIntended Audience :: System AdministratorsLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: PythonProgramming Language :: Python :: 2Programming Language :: Python :: 2.7Programming Language :: Python :: 3Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy |
Evidence: parquet-1.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “parquet file reader python”
- parquetReads Apache Parquet files in pure Python and outputs data as JSON or…
- fastparquetfastparquet reads and writes Apache Parquet files in Python, offering…
- parquet-toolsCommand-line tool to read, inspect, and export Parquet files from…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also fastparquet · gron · parquet-metadata · parquet-tools · pyexcel · sas7bdat · pyexcel-io · csv-diff · tabulate · tsv2py