scanpy
Single-Cell Analysis in Python.
Decision gist · record as of 2026-08-14
Yes. Scanpy is a mature, actively maintained toolkit with no known vulnerabilities, permissive licensing, and low install friction. It is the standard choice for single-cell gene expression analysis in Python. Install it if you work with single-cell RNA-seq or similar high-dimensional genomic data; the large dependency tree is a one-time cost for a comprehensive, well-documented analysis platform.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.12 or later.
- Large dependency tree (numpy, scipy, scikit-learn, pandas, matplotlib, etc.) may take time to resolve and install on first run.
- Low install friction with a pure-wheel distribution.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause (permissive) allows commercial and private use with minimal restrictions—you must include a copy of the license and the original copyright notice in distributions.
last release 2026-07-24 (21 days) · last repo commit 2026-08-14 · 2,542 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,038,260 downloads/mo, #4,459 on PyPI
Alternatives
Verify before relying
pip install scanpy
import scanpy as sc
adata = sc.read_h5ad('data.h5ad')
sc.pp.normalize_total(adata)
sc.tl.pca(adata)
sc.pl.pca(adata)- Whether dask integration for out-of-core analysis is production-ready or remains experimental as the description suggests.
- Performance characteristics and memory usage on datasets approaching or exceeding one million cells in practice.
What it is and what it does
Scanpy is a Python toolkit for single-cell genomics that handles the full workflow of analyzing gene expression data: reading and preprocessing raw counts, normalizing and scaling, dimensionality reduction, clustering, and statistical testing for differential expression. It is built on top of anndata for efficient data representation and integrates with the broader scverse ecosystem. The package is designed to scale from small pilot studies to datasets with over one million cells, and experimental dask support allows some operations on data larger than available memory.
The toolkit combines visualization (via matplotlib and seaborn), unsupervised learning (via scikit-learn and custom algorithms), and statistical inference (via statsmodels and scipy) into a unified API. It is actively maintained, production-stable, and widely used in academic and research settings for exploratory analysis, cell type discovery, and trajectory inference.
Use it for
- Preprocessing and quality control of raw single-cell RNA-seq count matrices before downstream analysis.
- Unsupervised clustering and visualization of cell populations to identify cell types or states.
- Differential expression testing between cell groups to find marker genes or disease-associated changes.
- Trajectory inference to reconstruct developmental or differentiation paths from snapshot data.
- Integration with external tools via anndata format for multi-tool pipelines in single-cell studies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Scanpy is a mature, actively maintained toolkit with no known vulnerabilities, permissive licensing, and low install friction. It is the standard choice for single-cell gene expression analysis in Python. Install it if you work with single-cell RNA-seq or similar high-dimensional genomic data; the large dependency tree is a one-time cost for a comprehensive, well-documented analysis platform.
Install
scanpy on PyPI
Before you install
Low install friction with a pure-wheel distribution. Active maintenance with a release 21 days ago and ongoing commits. Requires Python 3.12 or later. Brings 24 runtime dependencies including heavy scientific stacks (numpy, scipy, scikit-learn, pandas, matplotlib), so initial installation is substantial but straightforward.
Requires Python 3.12 or later. Large dependency tree (numpy, scipy, scikit-learn, pandas, matplotlib, etc.) may take time to resolve and install on first run.
License in practice
BSD-3-Clause (permissive) allows commercial and private use with minimal restrictions—you must include a copy of the license and the original copyright notice in distributions.
Quickstart
pip install scanpy
import scanpy as sc
adata = sc.read_h5ad('data.h5ad')
sc.pp.normalize_total(adata)
sc.tl.pca(adata)
sc.pl.pca(adata)
Verify before relying
- Whether dask integration for out-of-core analysis is production-ready or remains experimental as the description suggests.
- Performance characteristics and memory usage on datasets approaching or exceeding one million cells in practice.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.12 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 24 packagesanndatacertififast-array-utilsh5pyjobliblegacy-api-wrapmatplotlibnatsortnetworkxnumbanumpypackagingpandaspatsypynndescentscikit-learnscipyscverse-miscseabornsession-info2statsmodelstqdmtyping-extensionsumap-learn |
| Maintenance | Actively maintained 21 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,038,260 / month, #4,459 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleFramework :: JupyterIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/Engineering :: Bio-InformaticsTopic :: Scientific/Engineering :: Visualization |
Evidence: scanpy-1.12.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “single-cell gene expression analysis”
- scanpyScanpy is a toolkit for preprocessing, visualizing, clustering, and…
- cellxgene-censusProvides a Python API to query and retrieve cell metadata and…
- pydeseq2PyDESeq2 performs differential expression analysis on bulk RNA-seq…
Give your agent the search over MCP, or paste the wish link into any chat.
More Visualization packages
matplotlib creates static, animated, and interactive visualizations in Python, producing publication-quality figures in multiple formats for scripts, shells, web servers, and graphical interfaces.
Install it if you need to visualize data, generate publication-quality figures, or embed plots in applications.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Plotly is an interactive, browser-based graphing library that creates charts and visualizations from Python, rendering them as HTML that can be viewed in Jupyter notebooks, standalone files, or web applications.
Generates DOT language source code for graph structures and renders them using the Graphviz graph drawing software installed on your system.
Install it if you need to generate or render graphs from Python.
Streamlit transforms Python scripts into interactive web applications with minimal code, enabling rapid development of data dashboards, reports, and chat interfaces without requiring web development expertise.
Leather is a lightweight Python charting library for quick, no-frills data visualization. It generates charts without requiring perfect styling or extensive configuration.
See also pydeseq2 · scvi-tools · cobra · somacore · gseapy · cellxgene-census · spatialdata · tiledbsoma · anndata · peppy