pandavro
The interface between Avro and pandas DataFrame
Decision gist · record as of 2026-08-14
Yes, if you need to serialize pandas DataFrames to Avro or read Avro files into pandas. Install friction is low and the MIT license is permissive. Maintenance is aging (352 days since last release), but the repository is active and there are no known vulnerabilities. Be aware of the datetime timezone gotcha and the limitations on complex types and schema nesting before committing to it for a large project.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >= 3.9.0.
- Naive datetime columns must not contain timezone information; always use timezone-aware timestamps to avoid roundtrip errors due to fastavro's system timezone interpretation.
- Low install friction with a pure-Python wheel.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute pandavro freely provided you include the license notice.
last release 2025-08-27 (352 days) · last repo commit 2025-08-27 · 137 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 272,012 downloads/mo, #8,214 on PyPI
Alternatives
Verify before relying
pip install pandavro
import pandavro as pdx
import pandas as pd
df = pd.DataFrame({'col': [1, 2, 3]})
pdx.to_avro('file.avro', df)
df_read = pdx.read_avro('file.avro')- Whether the 352-day release gap reflects stable maintenance or reduced active development.
- Performance characteristics when handling large DataFrames or complex nested schemas.
- Compatibility with recent pandas and numpy versions beyond the stated support window.
What it is and what it does
Pandavro bridges Apache Avro and pandas, letting you serialize DataFrames to Avro files and deserialize Avro files back into DataFrames. It depends on fastavro for the low-level Avro I/O, pandas for the DataFrame abstraction, and numpy for type handling. The package infers Avro schemas from DataFrame dtypes, mapping numpy and pandas types (booleans, integers, floats, strings, timestamps, and complex types like records and arrays) to their Avro equivalents. All columns are treated as nullable by default, and nested records and arrays are supported.
The package handles most common data types but has known limitations: it does not support Avro enums, maps, or unions; it cannot infer non-nested schemas with DataFrame indexes; and naive datetime columns will not roundtrip correctly without explicit timezone information due to fastavro's design. You can provide a custom schema instead of relying on inference, and you can load nullable pandas datatypes (Int64Dtype, StringDtype, etc.) deterministically by passing na_dtypes=True to read_avro.
Use it for
- Export pandas DataFrames to Avro format for storage or transmission in data pipelines.
- Load Avro files from external sources (e.g., Kafka, data lakes) into pandas for analysis.
- Convert between pandas and Avro schemas when integrating with systems that use Avro as a standard format.
- Preserve nullable integer and string types when round-tripping DataFrames through Avro serialization.
- Infer Avro schemas automatically from DataFrame structure without manual schema definition.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to serialize pandas DataFrames to Avro or read Avro files into pandas.
Install friction is low and the MIT license is permissive. Maintenance is aging (352 days since last release), but the repository is active and there are no known vulnerabilities. Be aware of the datetime timezone gotcha and the limitations on complex types and schema nesting before committing to it for a large project.
Install
pandavro on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance status is aging—last release was 352 days ago—but the repository remains active and unarchived with recent commits.
Requires Python >= 3.9.0. Naive datetime columns must not contain timezone information; always use timezone-aware timestamps to avoid roundtrip errors due to fastavro's system timezone interpretation.
License in practice
MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute pandavro freely provided you include the license notice.
Quickstart
pip install pandavro
import pandavro as pdx
import pandas as pd
df = pd.DataFrame({'col': [1, 2, 3]})
pdx.to_avro('file.avro', df)
df_read = pdx.read_avro('file.avro')
Verify before relying
- Whether the 352-day release gap reflects stable maintenance or reduced active development.
- Performance characteristics when handling large DataFrames or complex nested schemas.
- Compatibility with recent pandas and numpy versions beyond the stated support window.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9.0 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesfastavropandasnumpy |
| Maintenance | Aging 352 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 272,012 / month, #8,214 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: pandavro-1.9.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “avro pandas dataframe”
- pandavroReads and writes Apache Avro files to and from pandas DataFrames,…
- hdfsHdfsCLI provides Python bindings and a command-line interface for…
- gspread-dataframeConverts between Google Sheets worksheets and pandas DataFrames,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also fastavro · avro-gen3 · avro-gen · pandas-read-xml · dbf · sklearn-pandas · gspread-dataframe · confluent_avro · sas7bdat · pyreadr