$npx skillfedfor your agent

kerchunk

Functions to make reference descriptions for ReferenceFileSystem

Worth itPyPI Scientific/EngineeringReleased Mar 2026138.0K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — kerchunk-0.2.10-py3-none-any.whl
v0.2.10 · released 2026-03-30 · Python >=3.11 · 5 runtime deps: fsspec, numcodecs, numpy, ujson, zarr

Yes. Kerchunk solves a specific, high-value problem for scientific and climate data workflows: efficient cloud access to legacy archival formats without data duplication. Low install friction, active maintenance, permissive licensing, no known vulnerabilities, and a focused dependency set make it a safe choice for teams working with large multi-file datasets in cloud environments.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.11 or later.
  • Source data files must be in a supported format (NetCDF, HDF5, GRIB, TIFF, FITS, or Zarr) and accessible via fsspec-supported storage backends.
  • Low install friction with a pure-Python wheel.

License · maintenance · safety

MIT (permissive) — MIT license (permissive) means you can use, modify, and distribute kerchunk freely in commercial and private projects with minimal restrictions, provided you include the license notice.

last release 2026-03-30 (137 days) · last repo commit 2026-03-30 · 367 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 137,953 downloads/mo, #11,348 on PyPI

Verify before relying

pip install kerchunk

import kerchunk.hdf
from fsspec.implementations.reference import ReferenceFileSystem

# Extract metadata from HDF5 file
reference_dict = kerchunk.hdf.SingleHdf5ToZarr('data.h5').translate()

# Access via virtual dataset
fs = ReferenceFileSystem(reference_dict)
data = fs.open('data', 'rb')
  • Specific performance gains or latency amortization metrics for concurrent chunk fetching compared to direct file access.
  • Scalability limits for datasets with millions of files and typical query response times.
  • Compatibility matrix for heterogeneous file types within a single virtual dataset.
  • Memory overhead of metadata consolidation for large multi-file datasets.
Same gist for agents: .md · .json

What it is and what it does

Kerchunk is a metadata extraction library that transforms archival data formats into cloud-friendly virtual datasets. Instead of copying or converting NetCDF, HDF5, GRIB, TIFF, FITS, or Zarr files, it extracts byte ranges, compression metadata, and structural information, storing this as a separate reference object. This allows you to create unified, queryable datasets spanning many source files without moving the original data.

The library integrates with fsspec to read from diverse storage backends—S3, GCS, HTTP, local filesystems, and network protocols—and with zarr for parallel, lock-free access. It supports asynchronous concurrent fetching of data chunks and coordinate-based indexing across arbitrary dimensions, making it a gateway for serverless, in-situ processing of massive archival datasets in the cloud while data providers continue using legacy formats.

Use it for

  • Aggregate NetCDF climate or weather data from thousands of files into a single queryable dataset without copying.
  • Access HDF5 scientific data stored in cloud object storage (S3, GCS) with efficient byte-range requests.
  • Create virtual GRIB datasets for meteorological analysis spanning multiple time steps and sources.
  • Build logical views over heterogeneous file types (mix of NetCDF and HDF5) with unified coordinate indexing.
  • Enable serverless data processing pipelines that fetch only required chunks from archival storage.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Kerchunk solves a specific, high-value problem for scientific and climate data workflows: efficient cloud access to legacy archival formats without data duplication. Low install friction, active maintenance, permissive licensing, no known vulnerabilities, and a focused dependency set make it a safe choice for teams working with large multi-file datasets in cloud environments.

Install

kerchunk on PyPI

Before you install

Low install friction with a pure-Python wheel. Active maintenance as of 2026-03-30 with 367 repository stars. Requires Python 3.11 or later. Five runtime dependencies (fsspec, numcodecs, numpy, ujson, zarr) are all widely used data-science libraries.

Requires Python 3.11 or later. Source data files must be in a supported format (NetCDF, HDF5, GRIB, TIFF, FITS, or Zarr) and accessible via fsspec-supported storage backends.

License in practice

MIT license (permissive) means you can use, modify, and distribute kerchunk freely in commercial and private projects with minimal restrictions, provided you include the license notice.

Quickstart

pip install kerchunk

import kerchunk.hdf
from fsspec.implementations.reference import ReferenceFileSystem

# Extract metadata from HDF5 file
reference_dict = kerchunk.hdf.SingleHdf5ToZarr('data.h5').translate()

# Access via virtual dataset
fs = ReferenceFileSystem(reference_dict)
data = fs.open('data', 'rb')

Verify before relying

  • Specific performance gains or latency amortization metrics for concurrent chunk fetching compared to direct file access.
  • Scalability limits for datasets with millions of files and typical query response times.
  • Compatibility matrix for heterogeneous file types within a single virtual dataset.
  • Memory overhead of metadata consolidation for large multi-file datasets.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.11
Install frictionLow. Pure-Python wheel
Runtime dependencies
5 packages
fsspecnumcodecsnumpyujsonzarr
MaintenanceActively maintained 137 days since the last release
Last repo commit
First released
Downloads137,953 / month, #11,348 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: PythonTopic :: Scientific/Engineering

Evidence: kerchunk-0.2.10-py3-none-any.whl

Tags

Capabilities
cloud-friendly data accessreference file system metadatavirtual datasets from multiple fileschunked data format accessin-situ cloud data processingmetadata extraction for archival formatsserverless data aggregation
Topics
cloud-storagescientific-datametadata-extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “cloud-friendly data access”

  • kerchunkKerchunk extracts metadata from chunked, compressed data formats…
  • access-parserParses Microsoft Access database files (.mdb and .accdb) in pure…
  • mo-dotsProvides null-safe dot-notation access to nested dictionaries and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also icechunk · earthkit-data · tensorstore · h5py · hickle · h5netcdf · cfgrib · netCDF4 · tables · datasets