pyensembl
Python interface to Ensembl reference genome metadata
Decision gist · record as of 2026-08-14
Yes. PyEnsembl is actively maintained, has low install friction, carries no known vulnerabilities, and is the standard tool for programmatic access to Ensembl genome metadata in Python. Install it if you need to query genes, transcripts, or exons by name or genomic position in a bioinformatics workflow. The main gotcha is the separate data-download step required before first use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later.
- Before first use, run `pyensembl install --release <N> --species <name>` to download and index genome data.
- Low friction: pure Python wheel with 7 runtime dependencies (datacache, gtfparse, numpy, and others).
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include a copy of the license and state any changes made.
last release 2026-07-08 (37 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 104,221 downloads/mo, #12,761 on PyPI
Alternatives
Verify before relying
pip install pyensembl
from pyensembl import EnsemblRelease
data = EnsemblRelease(77)
gene_names = data.gene_names_at_locus(contig=6, position=29945884)- Whether datacache, gtfparse, and other dependencies are actively maintained and free of known vulnerabilities.
- Performance characteristics when querying large genomes or handling non-Ensembl GTF formats (documentation notes this is still in development).
What it is and what it does
PyEnsembl is a Python library that wraps Ensembl reference genome data—exons, transcripts, genes, and their chromosomal locations—making it queryable through a Python API. It downloads GTF (gene transfer format) and FASTA sequence files from the Ensembl FTP server, indexes them into a local database, and exposes methods to look up genes by name or ID, find transcripts, retrieve exon coordinates, and query genomic features at specific loci. It also supports custom genomes loaded from user-supplied GTF and FASTA files or remote URLs.
The package is designed for bioinformatics workflows where you need to resolve gene names to genomic coordinates, retrieve transcript structures, or annotate genomic variants against a reference. It handles the download and caching automatically and stores data in a platform-specific cache directory (configurable via environment variable). The API provides many query methods—by gene ID, gene name, transcript ID, exon ID, chromosomal position, and strand—returning Gene, Transcript, and Exon objects with associated metadata.
Use it for
- Annotate genomic variants (SNPs, indels) by looking up overlapping genes and transcripts at their genomic coordinates.
- Retrieve transcript structures and exon boundaries for a given gene name to design primers or analyze splicing.
- Build a local reference database for a specific Ensembl release and species to support reproducible bioinformatics pipelines.
- Query the nearest gene to a genomic position when no gene directly overlaps, useful for intergenic variant classification.
- Load and index custom genome annotations (non-Ensembl GTF/FASTA) for organisms or assemblies not in the Ensembl repository.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
PyEnsembl is actively maintained, has low install friction, carries no known vulnerabilities, and is the standard tool for programmatic access to Ensembl genome metadata in Python. Install it if you need to query genes, transcripts, or exons by name or genomic position in a bioinformatics workflow. The main gotcha is the separate data-download step required before first use.
Install
pyensembl on PyPI
Before you install
Low friction: pure Python wheel with 7 runtime dependencies (datacache, gtfparse, numpy, and others). Marked active with a release 37 days ago. Requires Python 3.9 or later and a separate download step (`pyensembl install`) to fetch genome data before first use.
Requires Python 3.9 or later. Before first use, run `pyensembl install --release <N> --species <name>` to download and index genome data.
License in practice
Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include a copy of the license and state any changes made.
Quickstart
pip install pyensembl
from pyensembl import EnsemblRelease
data = EnsemblRelease(77)
gene_names = data.gene_names_at_locus(contig=6, position=29945884)
Verify before relying
- Whether datacache, gtfparse, and other dependencies are actively maintained and free of known vulnerabilities.
- Performance characteristics when querying large genomes or handling non-Ensembl GTF formats (documentation notes this is still in development).
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagestypechecksdatacachememoized-propertytinytimergtfparseserializablenumpy |
| Maintenance | Actively maintained 37 days since the last release |
| First released | |
| Downloads | 104,221 / month, #12,761 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Bio-Informatics |
Evidence: pyensembl-2.10.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “ensembl genome reference data python”
- pyensemblPyEnsembl provides a Python interface to query Ensembl reference…
- biocommons.seqrepoSeqRepo stores and retrieves biological sequences in a non-redundant,…
- refgenconfManages standardized configuration for reference genome assemblies,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Bio-Informatics packages
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
PubChemPy is a Python wrapper around the PubChem REST API that lets you search for chemical compounds by name, substructure, or similarity, retrieve their properties, and convert between chemical file formats.
Install it if you need programmatic access to PubChem data.
See also biocommons.seqrepo · gtfparse · bio · gprofiler-official · pyfaidx · pysam · varcode · mygene · pyranges · pybedtools