anndata
Annotated data.
Decision gist · record as of 2026-08-14
Yes. anndata is actively maintained, has no known vulnerabilities, low install friction, and fills a genuine need for annotated matrix storage in scientific Python. Its BSD-3-Clause permissive license poses no restriction. Install it if you work with structured scientific data—especially single-cell omics—or need sparse matrix support with metadata. The Python 3.12 or later requirement is a minor constraint for modern projects.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.12 or later.
- Low install friction; pure Python wheel with 11 runtime dependencies including numpy, pandas, scipy, and zarr.
- Active maintenance with a release 32 days ago and ongoing commits.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause permissive license allows commercial and private use with minimal restrictions—standard for scientific Python packages.
last release 2026-07-13 (32 days) · last repo commit 2026-08-14 · 764 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,812,748 downloads/mo, #3,528 on PyPI
Alternatives
Verify before relying
pip install anndata
import anndata
import numpy as np
X = np.random.randn(10, 5)
adata = anndata.AnnData(X=X)
adata.write_h5ad('data.h5ad')- Performance characteristics and scalability limits for terabyte-scale datasets mentioned in description.
- Specific lazy-loading behavior and memory footprint compared to alternatives.
- Integration depth with annbatch for minibatch data loading workflows.
What it is and what it does
anndata is a Python library for storing and manipulating annotated data matrices—arrays with associated metadata, observations, and variables. It sits between pandas DataFrames and xarray Datasets, designed to handle scientific data efficiently whether in memory or on disk. The package supports sparse matrices, lazy operations, and multiple storage backends including HDF5 and Zarr, making it suitable for datasets that range from small in-memory arrays to larger disk-backed collections.
The library was originally built for Scanpy, a single-cell analysis toolkit, but has become a general-purpose container for annotated matrices across bioinformatics and other scientific domains. It depends on numpy, pandas, scipy, h5py, zarr, and several utility packages, giving it a stable foundation in the scientific Python ecosystem. The package is actively maintained as part of the scverse project and fiscally sponsored by NumFOCUS.
Use it for
- Store and manipulate single-cell RNA-seq data with cell metadata and gene annotations.
- Work with large sparse matrices that don't fit efficiently in pandas or numpy alone.
- Persist annotated datasets to disk in HDF5 or Zarr format for reproducible analysis.
- Build data pipelines that combine in-memory and disk-backed operations with lazy loading.
- Integrate with Scanpy or other scverse tools for downstream bioinformatics workflows.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
anndata is actively maintained, has no known vulnerabilities, low install friction, and fills a genuine need for annotated matrix storage in scientific Python. Its BSD-3-Clause permissive license poses no restriction. Install it if you work with structured scientific data—especially single-cell omics—or need sparse matrix support with metadata. The Python 3.12 or later requirement is a minor constraint for modern projects.
Install
anndata on PyPI
Before you install
Low install friction; pure Python wheel with 11 runtime dependencies including numpy, pandas, scipy, and zarr. Active maintenance with a release 32 days ago and ongoing commits.
Requires Python 3.12 or later.
License in practice
BSD-3-Clause permissive license allows commercial and private use with minimal restrictions—standard for scientific Python packages.
Quickstart
pip install anndata
import anndata
import numpy as np
X = np.random.randn(10, 5)
adata = anndata.AnnData(X=X)
adata.write_h5ad('data.h5ad')
Verify before relying
- Performance characteristics and scalability limits for terabyte-scale datasets mentioned in description.
- Specific lazy-loading behavior and memory footprint compared to alternatives.
- Integration depth with annbatch for minibatch data loading workflows.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.12 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 11 packagesarray-api-compath5pylegacy-api-wrapnatsortnumpypackagingpandasscipyscverse-misctyping-extensionszarr |
| Maintenance | Actively maintained 32 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,812,748 / month, #3,528 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Environment :: ConsoleFramework :: JupyterIntended Audience :: DevelopersIntended Audience :: Science/ResearchNatural Language :: EnglishOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/Engineering :: Bio-InformaticsTopic :: Scientific/Engineering :: Visualization |
Evidence: anndata-0.13.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “annotated data matrix storage”
- anndataanndata handles annotated data matrices in memory and on disk with…
- synapse-s3-storage-providerA Synapse media storage provider that integrates Amazon S3 (or…
- kaldiioKaldiio reads and writes Kaldi archive (ark) and script (scp) files…
Give your agent the search over MCP, or paste the wish link into any chat.
More Visualization packages
matplotlib creates static, animated, and interactive visualizations in Python, producing publication-quality figures in multiple formats for scripts, shells, web servers, and graphical interfaces.
Install it if you need to visualize data, generate publication-quality figures, or embed plots in applications.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Plotly is an interactive, browser-based graphing library that creates charts and visualizations from Python, rendering them as HTML that can be viewed in Jupyter notebooks, standalone files, or web applications.
Generates DOT language source code for graph structures and renders them using the Graphviz graph drawing software installed on your system.
Install it if you need to generate or render graphs from Python.
Streamlit transforms Python scripts into interactive web applications with minimal code, enabling rapid development of data dashboards, reports, and chat interfaces without requiring web development expertise.
Leather is a lightweight Python charting library for quick, no-frills data visualization. It generates charts without requiring perfect styling or extensive configuration.
See also mudata · spatialdata · xarray · tabmat · scvi-tools · formulaic · sparse-dot-topn · linopy · spark-sklearn · scanpy