pgpq
Arrow -> PostgreSQL encoder
Decision gist · record as of 2026-08-14
Yes, if you regularly load Arrow data into PostgreSQL and want to avoid row-by-row inserts or intermediate file formats. The package is actively maintained, has no known vulnerabilities, and solves a specific ETL bottleneck. Medium install friction (compiled wheels) is typical for performance-critical libraries. Not necessary if you use PostgreSQL's native tools or rarely bulk-load Arrow data.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; PostgreSQL server and psycopg connection for actual data transfer.
- Medium install friction due to compiled wheels (cp39-abi3 binaries for multiple platforms).
- Active maintenance with recent commits and steady releases since early 2023.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; no copyleft obligations.
last release 2026-02-28 (167 days) · last repo commit 2026-08-14 · 280 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 173,253 downloads/mo, #10,314 on PyPI
Alternatives
Verify before relying
pip install pgpq
from pgpq import ArrowToPostgresBinaryEncoder
import pyarrow.dataset as ds
dataset = ds.dataset('/path/to/parquet')
encoder = ArrowToPostgresBinaryEncoder(dataset.schema)
pg_schema = encoder.schema()
# Use encoder.write_header(), encoder.write_batch(batch), encoder.finish() with psycopg COPY- Performance characteristics compared to alternative bulk-load methods or native PostgreSQL tooling.
- Support for all PostgreSQL data types or limitations in type mapping beyond the documented examples.
What it is and what it does
pgpq is a Python library that translates PyArrow RecordBatches into PostgreSQL's native binary wire format, enabling direct bulk loading of Arrow data into Postgres via the COPY command. It handles schema translation, field encoding, and binary serialization so you can stream Arrow datasets (from Parquet, CSV, or other sources) directly into Postgres without intermediate conversions or row-by-row inserts.
The package sits between PyArrow and psycopg, providing an ArrowToPostgresBinaryEncoder that takes an Arrow schema, optionally customizes field encoders (e.g., mapping string columns to JSONB), and produces the byte stream PostgreSQL expects. You define a temporary or permanent table matching the encoder's schema, then pipe the encoded batches through psycopg's COPY interface. It's designed for data warehouse and ETL workflows where bulk loading speed matters.
Use it for
- Load Parquet files from S3 or local disk into PostgreSQL without staging or intermediate formats.
- Bulk import Arrow datasets with custom field encoding (e.g., string-to-JSONB) to match Postgres table schemas.
- Stream large datasets through psycopg's COPY interface while avoiding row-by-row inserts or temporary CSV files.
- Automate data pipelines that extract Arrow data and load it into Postgres for analytics or reporting.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you regularly load Arrow data into PostgreSQL and want to avoid row-by-row inserts or intermediate file formats.
The package is actively maintained, has no known vulnerabilities, and solves a specific ETL bottleneck. Medium install friction (compiled wheels) is typical for performance-critical libraries. Not necessary if you use PostgreSQL's native tools or rarely bulk-load Arrow data.
Install
pgpq on PyPI
Before you install
Medium install friction due to compiled wheels (cp39-abi3 binaries for multiple platforms). Active maintenance with recent commits and steady releases since early 2023.
Requires Python 3.10 or later; PostgreSQL server and psycopg connection for actual data transfer.
License in practice
MIT license permits commercial and private use with minimal restrictions; no copyleft obligations.
Quickstart
pip install pgpq
from pgpq import ArrowToPostgresBinaryEncoder
import pyarrow.dataset as ds
dataset = ds.dataset('/path/to/parquet')
encoder = ArrowToPostgresBinaryEncoder(dataset.schema)
pg_schema = encoder.schema()
# Use encoder.write_header(), encoder.write_batch(batch), encoder.finish() with psycopg COPY
Verify before relying
- Performance characteristics compared to alternative bulk-load methods or native PostgreSQL tooling.
- Support for all PostgreSQL data types or limitations in type mapping beyond the documented examples.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packagepyarrow |
| Maintenance | Actively maintained 167 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 173,253 / month, #10,314 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: pgpq-0.11.1-cp39-abi3-macosx_10_12_x86_64.macosx_11_0_arm64.macosx_10_12_universal2.whl; pgpq-0.11.1-cp39-abi3-macosx_10_12_x86_64.whl; pgpq-0.11.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; pgpq-0.11.1-cp39-abi3-manylinux_2_5_i686.manylinux1_i686.whl; pgpq-0.11.1-cp39-abi3-musllinux_1_2_i686.whl; pgpq-0.11.1-cp39-abi3-musllinux_1_2_x86_64.whl; pgpq-0.11.1-cp39-abi3-win32.whl; pgpq-0.11.1-cp39-abi3-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “arrow to postgres binary encoder”
- pgpqEncodes PyArrow RecordBatches into PostgreSQL's native binary format…
- pgcopypgcopy wraps PostgreSQL's binary COPY protocol to load data into…
- py-ubjsonEncodes and decodes data in Universal Binary JSON format, a compact…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also pgcopy · pydantic-to-pyarrow · adbc-driver-postgresql · django-postgres-copy · pg0-embedded · pysqlsync · pyarrow · arrow-odbc · geoarrow-pyarrow · sling