pyfaidx
pyfaidx: efficient pythonic random access to fasta subsequences
Decision gist · record as of 2026-08-14
Yes. pyfaidx is production-stable, actively maintained, has no known vulnerabilities, and solves a real problem in bioinformatics workflows. Low install friction and permissive licensing make it a straightforward choice for anyone working with FASTA files in Python.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction: pure Python with only two lightweight runtime dependencies (importlib_metadata and packaging).
- Active maintenance with a recent release 148 days ago and 489 repository stars.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause is permissive; you can use this package freely in commercial and proprietary projects with minimal restrictions.
last release 2026-03-19 (148 days) · last repo commit 2026-03-19 · 489 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 529,907 downloads/mo, #6,159 on PyPI
Alternatives
Verify before relying
pip install pyfaidx
from pyfaidx import Fasta
genes = Fasta('genes.fasta')
seq = genes['sequence_id'][200:230]
print(seq.seq)- Whether the command-line faidx script is reliably available after installation on all supported platforms.
- Performance characteristics when working with very large FASTA files or many concurrent accesses.
What it is and what it does
pyfaidx is a pure Python implementation of samtools' faidx functionality, enabling efficient random access to any subsequence in a FASTA file without loading the entire file into memory. It creates a small flat index file (.fai) that allows seeking directly to the sequence you need. The package provides both a Python API (with dictionary-like and method-based access) and a command-line tool for FASTA manipulation without programming.
The library supports slicing, reverse complements, spliced sequences, and coordinate transformations (1-based and 0-based). It works with sequences indexed by name or position, handles custom key functions for flexible naming schemes, and can filter sequences during indexing. The API is compatible with pygr's seqdb module, making it a drop-in replacement for existing bioinformatics workflows.
Use it for
- Extract specific genomic regions from reference genomes for variant analysis or annotation pipelines.
- Build sequence databases for rapid lookup of genes or transcripts by identifier without full file loads.
- Perform reverse-complement operations on DNA sequences for primer design or alignment validation.
- Slice and manipulate FASTA sequences programmatically in Python-based bioinformatics workflows.
- Use the faidx command-line tool to extract or modify FASTA records without writing custom scripts.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
pyfaidx is production-stable, actively maintained, has no known vulnerabilities, and solves a real problem in bioinformatics workflows. Low install friction and permissive licensing make it a straightforward choice for anyone working with FASTA files in Python.
Install
pyfaidx on PyPI
Before you install
Low friction: pure Python with only two lightweight runtime dependencies (importlib_metadata and packaging). Active maintenance with a recent release 148 days ago and 489 repository stars.
License in practice
BSD-3-Clause is permissive; you can use this package freely in commercial and proprietary projects with minimal restrictions.
Quickstart
pip install pyfaidx
from pyfaidx import Fasta
genes = Fasta('genes.fasta')
seq = genes['sequence_id'][200:230]
print(seq.seq)
Verify before relying
- Whether the command-line faidx script is reliably available after installation on all supported platforms.
- Performance characteristics when working with very large FASTA files or many concurrent accesses.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesimportlib_metadatapackaging |
| Maintenance | Actively maintained 148 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 529,907 / month, #6,159 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: UnixProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: PyPyTopic :: Scientific/Engineering :: Bio-Informatics |
Evidence: pyfaidx-0.9.0.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “fasta file random access”
- pyfaidxProvides fast random access to subsequences in FASTA files using a…
- pysampysam reads, manipulates, and writes genomic data files…
- dnaiodnaio reads and writes FASTQ, FASTA, and uBAM files with optimized…
Give your agent the search over MCP, or paste the wish link into any chat.
More Bio-Informatics packages
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
PubChemPy is a Python wrapper around the PubChem REST API that lets you search for chemical compounds by name, substructure, or similarity, retrieve their properties, and convert between chemical file formats.
Install it if you need programmatic access to PubChem data.
See also pysam · pyensembl · dnaio · biocommons.seqrepo · cyvcf2 · fuzzysearch · pyranges · pyteomics · bx-python · RUST