pydeseq2
A python implementation of DESeq2.
Decision gist · record as of 2026-08-14
Yes, if you work with bulk RNA-seq data in Python and need DESeq2-compatible statistical analysis. The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install with caution if you require exact numerical equivalence to the original R DESeq2 or advanced features beyond Wald tests—verify feature completeness against your specific experimental design first.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.11 or higher; input data must be in AnnData format or convertible to it.
- Low install friction with a pure-wheel distribution.
- Actively maintained as of August 2026 with recent commits; transferred to scverse community maintenance in December 2025.
License · maintenance · safety
permissive license (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for academic, commercial, and proprietary projects.
last release 2026-01-23 (203 days) · last repo commit 2026-08-10 · 759 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 178,023 downloads/mo, #10,202 on PyPI
Alternatives
Verify before relying
pip install pydeseq2
import anndata
from pydeseq2.dds import DeseqDataSet
# Load count data into AnnData object
ads = anndata.read_h5ad('counts.h5ad')
# Initialize and run DESeq2 analysis
dds = DeseqDataSet(adata=ads, design_factors=['condition'])
dds.deseq2()- Completeness of feature parity with DESeq2 v1.34.0 beyond single/multi-factor Wald tests
- Performance characteristics on large-scale datasets (sample size, gene count thresholds)
- Availability of downstream analysis tools (e.g., visualization, result filtering) within the package
What it is and what it does
PyDESeq2 is a Python port of the widely-used R package DESeq2, designed to make differential expression analysis accessible to Python users working with bulk RNA-seq data. It implements statistical methods for comparing gene expression across experimental conditions, handling count normalization, dispersion estimation, and hypothesis testing. The package works with data in AnnData format and depends on standard scientific Python libraries: numpy, pandas, scipy, scikit-learn, matplotlib, and formulaic for design matrix specification.
The implementation currently covers single-factor and multi-factor experimental designs using Wald tests, matching the default behavior of DESeq2 v1.34.0. As a re-implementation from scratch rather than a direct wrapper, it may produce slightly different numerical results or lack some advanced features of the original R package. The project is actively maintained by the scverse community and welcomes feature requests via issue tracking.
Use it for
- Compare gene expression between treatment and control groups in bulk RNA-seq experiments
- Analyze multi-factor designs (e.g., treatment × genotype) to identify condition-specific effects
- Normalize and standardize RNA-seq count data for downstream statistical testing
- Integrate differential expression results into Python-based bioinformatics pipelines using AnnData
- Identify significantly differentially expressed genes with multiple testing correction
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with bulk RNA-seq data in Python and need DESeq2-compatible statistical analysis.
The package has low install friction, active maintenance, permissive licensing, and no known vulnerabilities. Install with caution if you require exact numerical equivalence to the original R DESeq2 or advanced features beyond Wald tests—verify feature completeness against your specific experimental design first.
Install
pydeseq2 on PyPI
Before you install
Low install friction with a pure-wheel distribution. Actively maintained as of August 2026 with recent commits; transferred to scverse community maintenance in December 2025. Tested against Python 3.11–3.13 with current versions of its eight runtime dependencies.
Requires Python 3.11 or higher; input data must be in AnnData format or convertible to it.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for academic, commercial, and proprietary projects.
Quickstart
pip install pydeseq2
import anndata
from pydeseq2.dds import DeseqDataSet
# Load count data into AnnData object
ads = anndata.read_h5ad('counts.h5ad')
# Initialize and run DESeq2 analysis
dds = DeseqDataSet(adata=ads, design_factors=['condition'])
dds.deseq2()
Verify before relying
- Completeness of feature parity with DESeq2 v1.34.0 beyond single/multi-factor Wald tests
- Performance characteristics on large-scale datasets (sample size, gene count thresholds)
- Availability of downstream analysis tools (e.g., visualization, result filtering) within the package
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagesanndataformulaic-contrastsformulaicmatplotlibnumpypandasscikit-learnscipy |
| Maintenance | Actively maintained 203 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 178,023 / month, #10,202 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13 |
Evidence: pydeseq2-0.5.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “differential expression analysis RNA-seq”
- pydeseq2PyDESeq2 performs differential expression analysis on bulk RNA-seq…
- scanpyScanpy is a toolkit for preprocessing, visualizing, clustering, and…
- gseapyPerforms gene set enrichment analysis (GSEA) on genomic data using…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also scanpy · ViennaRNA · multiqc · gseapy · cellxgene-census · gtfparse · pyhmmer · formulaic-contrasts · cobra · pyensembl