Packages
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
PubChemPy is a Python wrapper around the PubChem REST API that lets you search for chemical compounds by name, substructure, or similarity, retrieve their properties, and convert between chemical file formats.
Install it if you need programmatic access to PubChem data.
SimSIMD provides SIMD-optimized kernels for computing vector distances, dot-products, and similarity measures across multiple data types and precisions, with support for spatial, probabilistic, and bit-level operations.
Install it if vector similarity or distance computation is a measurable bottleneck in your application.
igraph provides a Python interface to a high-performance C graph library for constructing, analyzing, and visualizing networks and complex graphs.
Biotite provides a unified Python library for computational molecular biology, handling sequence and biomolecular structure data through file I/O, analysis, visualization, and database integration.
Reads and writes molecular dynamics trajectory files in XTC, TRR, NetCDF, and DCD formats, extracted from MDTraj as a standalone library for use by Biotite.
However, maintenance is aging (last release 650 days ago), so verify that the formats you need are fully supported and check Biotite's current requirements before…
anndata handles annotated data matrices in memory and on disk with sparse data support, lazy operations, and efficient storage—positioned as a middle ground between pandas and xarray for scientific data.
Install it if you work with structured scientific data—especially single-cell omics—or need sparse matrix support with metadata.
PyHMMER provides Python bindings to HMMER3, a biological sequence analysis tool that uses profile hidden Markov models to search for sequence homologs in protein or DNA databases.
A Python SDK for web scraping, crawling, searching, and extracting structured data from websites and research papers via the Firecrawl API, returning results as clean Markdown, HTML, or typed objects.
biothings_client provides Python wrappers to query gene, variant, chemical, disease, geneset, and taxon data from BioThings API services via synchronous or asynchronous clients.
Install it if you need programmatic access to gene, variant, chemical, disease, or taxon data; skip it only if you prefer direct HTTP calls or need a different data…
Mygene is a Python wrapper for the MyGene.Info REST API that retrieves and queries gene annotation data, including symbols, names, and cross-references across multiple species.
Python interface to g:Profiler toolkit for functional enrichment analysis, gene identifier conversion, and ortholog mapping across organisms.
However, the package is abandoned (last release 2019-04-02) and you should verify that the upstream g:Profiler service itself is still maintained and accessible…
Gemmi is a C++ library with Python bindings for reading, writing, and manipulating macromolecular structural data (PDB, mmCIF, MTZ files) and crystallographic information, including symmetries, density maps, and refinement restraints.
Scanpy is a toolkit for preprocessing, visualizing, clustering, and analyzing single-cell gene expression data, handling datasets from hundreds to millions of cells efficiently.
Install it if you work with single-cell RNA-seq or similar high-dimensional genomic data; the large dependency tree is a one-time cost for a comprehensive,…
pgmpy provides data structures and algorithms for causal discovery, causal inference, probabilistic inference, parameter learning, and model validation across Bayesian networks, DAGs, and structural equation models.
Tests whether a graph is planar, computes planar embeddings, draws planar graphs as ASCII art, and isolates forbidden subgraphs using algorithms from the Edge Addition Planarity Suite.
pybigtools is a Python wrapper around a Rust library for reading and writing bigWig and bigBed genomic data files with high performance and minimal memory overhead.
Install it if you work with genomic interval data in these formats.
Python interface to the igraph graph library for network analysis and complex graph algorithms; this is a legacy package superseded by the igraph package.
Provides fast random access to subsequences in FASTA files using a samtools-compatible index, with a pure Python implementation for indexing, retrieval, and in-place modification.
NumKong provides mixed-precision linear algebra and distance kernels with automatic accumulator widening, GIL-free batched operations, and low-precision dtype support (BFloat16, Float8, Float6, packed bits) across x86, ARM, RISC-V, and other architectures.
PROPKA predicts pKa values of ionizable groups in proteins and protein-ligand complexes from 3D structure using empirical heuristics.
However, dormant maintenance and lack of recent commits mean you should test predictions against your own data or literature benchmarks before using results in new…
primer3-py wraps the Primer3 library to provide fast oligo thermodynamic analysis (melting temperature, hairpin formation) and primer design through a Python API.
mrcfile reads and writes MRC2014 format files used in structural biology to store image and volume data, exposing headers and data as numpy arrays through a simple Python API.
Install it if you work with MRC format files in cryo-EM or related fields; skip it if you have no need for MRC file I/O.
PDB2PQR prepares protein structures from PDB files for biomolecular simulations and continuum solvation calculations by automating structure setup, protonation, and parameterization tasks.
Provides a pure Python interface to read, parse, and work with PDBx/mmCIF molecular structure data files from the Protein Data Bank.
Embeds an interactive 3D molecular viewer in Jupyter notebooks using 3Dmol.js, allowing visualization and manipulation of protein and molecular structures directly in the notebook.
alchemlyb parses molecular dynamics simulation output, extracts uncorrelated samples from timeseries data, and estimates free energies using MBAR, BAR, and thermodynamic integration methods.
Install it if you work with alchemical free energy calculations; skip it if you do not.
Vesin computes neighbor lists for atomistic systems—identifying which atoms are within a cutoff distance of each other—with a Python interface backed by compiled code for speed.
Retrieves and manages semantic prefix maps that convert compact URIs (like `skos`) to their full namespace URIs, with support for multiple sources and clash-resilient prefix ordering.
Converts between CURIEs (compact URI identifiers like GO:0008150) and full URIs using JSON-LD contexts.
Converts nested JSON/YAML objects into flat tabular rows suitable for dataframes, spreadsheets, or databases, with the ability to round-trip back to the original structure.
However, do not expect active development or bug fixes—evaluate whether the current feature set meets your needs before committing to a dependency.
PyCAT-Napari is a napari-based desktop application for detecting, measuring, and analyzing biomolecular condensates in fluorescence and brightfield microscopy images, with support for 2D, Z-stack, and time-series data.
cwltool is the reference implementation of the Common Workflow Language standard, enabling you to write, validate, and execute portable scientific workflows defined in CWL format across different computing environments.
Install it if you need to execute portable workflows, validate CWL definitions, or integrate CWL workflows into Python applications.
Pyteomics provides Python tools for proteomics data analysis, including mass calculation, sequence manipulation, and access to MS/LC-MS data, FASTA databases, and search engine outputs.
Install it if you work with mass spectrometry data, peptide sequences, or proteomics search results in Python.
OpenSlide Python provides a Python interface to read whole-slide images (virtual slides) used in digital pathology, supporting multiple vendor formats and efficient multi-resolution access to gigabyte-scale images.
The main gotcha is the separate OpenSlide C library dependency.
Computes neighbor lists for atomistic systems efficiently, with support for periodic boundary conditions and GPU execution via PyTorch.