$npx skillfedfor your agent

pyarrowfs-adlgen2

Use pyarrow with Azure Data Lake gen2

With conditionsPyPI DatabaseReleased Jun 2024116.0K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pyarrowfs_adlgen2-0.2.5-py3-none-any.whl
v0.2.5 · released 2024-06-27 · Python >=3.6 · 2 runtime deps: pyarrow, azure-storage-file-datalake

Yes, if you need to read or write Parquet on Azure Data Lake Gen2 with PyArrow. Install friction is low, dependencies are stable, and the MIT license is unrestricted. The dormant maintenance status is not a risk—the package is explicitly stable with no major features planned. The main caveat is that dormancy means no active bug fixes or feature development, so evaluate whether your use case aligns with the current feature set.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Azure credentials configured (e.g., via `az login` or environment variables) for azure.identity.DefaultAzureCredential() to work.
  • Low friction install with two stable runtime dependencies (pyarrow and azure-storage-file-datalake).
  • Maintenance is dormant—last commit was 2024-06-27 and no releases in 778 days—but the package is explicitly described as stable with no major features planned, so dormancy reflects maturity rather than abandonment.

License · maintenance · safety

MIT (permissive) — MIT license (permissive) imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

last release 2024-06-27 (778 days) · last repo commit 2024-06-27 · 29 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 115,982 downloads/mo, #12,225 on PyPI

Verify before relying

pip install pyarrowfs-adlgen2

import azure.identity
import pyarrow.fs
import pyarrowfs_adlgen2

handler = pyarrowfs_adlgen2.AccountHandler.from_account_name(
    'YOUR_ACCOUNT_NAME', azure.identity.DefaultAzureCredential())
fs = pyarrow.fs.PyFileSystem(handler)
ds = pyarrow.dataset.dataset('container/dataset.parq', filesystem=fs)
table = ds.to_table()
  • Whether the performance advantage over other filesystem implementations holds across different dataset sizes and structures beyond the NYC taxi benchmark.
  • Compatibility with the latest versions of pyarrow and azure-storage-file-datalake beyond what classifiers declare.
Same gist for agents: .md · .json

What it is and what it does

pyarrowfs-adlgen2 bridges PyArrow and Azure Data Lake Gen2 by implementing a PyArrow filesystem that lets you read and write Parquet files directly from cloud storage. Instead of downloading data locally first, you pass the filesystem handler to PyArrow's dataset API and work with remote files as if they were local. It wraps the azure-storage-file-datalake SDK, which provides fast directory listing.

The package is small and stable, supporting Python 3.6 through 3.11. You authenticate via azure.identity, configure optional timeouts, and then use standard PyArrow patterns: read operations with a filesystem argument, or write datasets for PyArrow 3 or greater. The package is dormant but intentionally so—it has a minimal API and no planned major features.

Use it for

  • Read multi-file Parquet datasets from Azure Data Lake Gen2 into PyArrow tables without downloading to local disk.
  • Stream large Parquet datasets from Azure for analysis or transformation in memory using PyArrow's columnar operations.
  • Write partitioned Parquet datasets back to Azure Data Lake Gen2 directly from PyArrow tables.
  • Integrate Azure Data Lake Gen2 as a data source in ETL pipelines that already use PyArrow.
  • Access a single container or filesystem within an Azure storage account when full account access is not needed or desired.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to read or write Parquet on Azure Data Lake Gen2 with PyArrow.

Install friction is low, dependencies are stable, and the MIT license is unrestricted. The dormant maintenance status is not a risk—the package is explicitly stable with no major features planned. The main caveat is that dormancy means no active bug fixes or feature development, so evaluate whether your use case aligns with the current feature set.

Install

pyarrowfs-adlgen2 on PyPI

Before you install

Low friction install with two stable runtime dependencies (pyarrow and azure-storage-file-datalake). Maintenance is dormant—last commit was 2024-06-27 and no releases in 778 days—but the package is explicitly described as stable with no major features planned, so dormancy reflects maturity rather than abandonment.

Requires Azure credentials configured (e.g., via `az login` or environment variables) for azure.identity.DefaultAzureCredential() to work.

License in practice

MIT license (permissive) imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.

Quickstart

pip install pyarrowfs-adlgen2

import azure.identity
import pyarrow.fs
import pyarrowfs_adlgen2

handler = pyarrowfs_adlgen2.AccountHandler.from_account_name(
    'YOUR_ACCOUNT_NAME', azure.identity.DefaultAzureCredential())
fs = pyarrow.fs.PyFileSystem(handler)
ds = pyarrow.dataset.dataset('container/dataset.parq', filesystem=fs)
table = ds.to_table()

Verify before relying

  • Whether the performance advantage over other filesystem implementations holds across different dataset sizes and structures beyond the NYC taxi benchmark.
  • Compatibility with the latest versions of pyarrow and azure-storage-file-datalake beyond what classifiers declare.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.6
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
pyarrowazure-storage-file-datalake
MaintenanceDormant 778 days since the last release
Last repo commit
First released
Downloads115,982 / month, #12,225 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: pyarrowfs_adlgen2-0.2.5-py3-none-any.whl

Tags

Capabilities
azure data lake gen2 pyarrowparquet azure datalake filesystemread parquet from azurepyarrow azure storageazure datalake gen2 connector
Topics
azure-cloudparquetdata-lake
PyPI keywords
azuredatalakefilesystempyarrowparquet

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “azure data lake gen2 pyarrow”

  • pyarrowfs-adlgen2Provides a PyArrow filesystem interface for reading and writing…
  • adlfsProvides a filesystem interface to Azure Blob Storage and Azure Data…
  • azure-storage-file-datalakeProvides Python client library for Azure Data Lake Storage Gen2,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Database packages

psycopg2-binary Worth it
PyPI · Software Development · released Apr 2026

psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.

copyleftcompiled wheel · 3.9+
271.6Mdownloads / mo
redis Worth it
PyPI · Database · released Jul 2026

Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.

Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.

MITpure Python · 3.10+
268.3Mdownloads / mo
ydb Worth it
PyPI · Database · released Jul 2026

YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.

Install it if you need to connect Python applications to YDB databases.

permissive licensepure Python · 3.10+
210.0Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
sqlparse Worth it
PyPI · Software Development · released Aug 2026

sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.

Install it if you need to manipulate, format, or analyze SQL text programmatically.

BSD-3-Clausepure Python · 3.10+
148.9Mdownloads / mo
dbt-adapters With conditions
PyPI · Database · released Jul 2026

Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.

Apache-2.0pure Python · 3.10.0+
121.3Mdownloads / mo

See also adlfs · azure-storage-file-datalake · azure-datalake-store · delta-sharing · fastparquet · pyarrow-hotfix · hops-deltalake · pydantic-to-pyarrow · geoarrow-pyarrow · azure-storage-blob