dnaio
Read and write FASTA and FASTQ files efficiently
Decision gist · record as of 2026-08-14
Yes, if you work with FASTQ, FASTA, or uBAM files in bioinformatics. dnaio is production-stable (Development Status 5), has no known vulnerabilities, and offers efficient parsing with minimal dependencies (only xopen). The MIT license is unrestricted. Maintenance is aging (306 days since last release), but the repo remains active and the library is mature. Install friction is medium due to compiled components, but prebuilt wheels are available for modern Python and common platforms.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Optional: install zstandard for .zst file support.
- Medium install friction due to compiled components (Cython-based), but prebuilt wheels cover modern Python versions (3.10–3.14) and major platforms.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute dnaio freely in commercial and private projects with minimal restrictions.
last release 2025-10-12 (306 days) · last repo commit 2025-10-12 · 70 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 92,387 downloads/mo, #13,454 on PyPI
Alternatives
Verify before relying
pip install dnaio
import dnaio
with dnaio.open("reads.fastq.gz") as f:
for record in f:
print(record.name, len(record.sequence))- Performance benchmarks comparing dnaio to other FASTQ parsers in the ecosystem
- Memory footprint for large files (e.g., multi-gigabyte FASTQ datasets)
- Interleaved paired-end parsing behavior and API details
What it is and what it does
dnaio is a Python library for reading and writing FASTQ, FASTA, and uBAM sequence files with a focus on parsing speed and efficiency. It originated as part of the Cutadapt tool and has been refined since becoming standalone. The library automatically detects and handles compressed files (.gz, .bz2, .xz, .zst) and supports both single-file and two-file paired-end workflows, as well as interleaved paired-end formats. It can parse uBAM files directly from ONT basecallers like dorado.
The main entry point is dnaio.open(), which returns an iterator over sequence records. Each record exposes name, sequence, and quality attributes. The library prioritizes FASTQ and uBAM parsing; FASTA parsing is supported but not as heavily optimized. It does not support multi-line FASTQ files. Installation is straightforward via pip, with optional zstandard support for .zst compression.
Use it for
- Parse large FASTQ files from sequencing runs in genomics pipelines without loading entire files into memory
- Read uBAM output directly from ONT dorado basecaller for real-time or batch processing
- Handle automatically-compressed FASTQ/FASTA files across multiple compression formats in a single workflow
- Process paired-end sequencing data stored in two separate files or a single interleaved file
- Extract sequence and quality information for downstream bioinformatics analysis (alignment, QC, filtering)
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with FASTQ, FASTA, or uBAM files in bioinformatics.
dnaio is production-stable (Development Status 5), has no known vulnerabilities, and offers efficient parsing with minimal dependencies (only xopen). The MIT license is unrestricted. Maintenance is aging (306 days since last release), but the repo remains active and the library is mature. Install friction is medium due to compiled components, but prebuilt wheels are available for modern Python and common platforms.
Install
dnaio on PyPI
Before you install
Medium install friction due to compiled components (Cython-based), but prebuilt wheels cover modern Python versions (3.10–3.14) and major platforms. Last release was 306 days ago; repo is active and not archived, though maintenance pace is aging.
Requires Python 3.10 or later. Optional: install zstandard for .zst file support.
License in practice
MIT license is permissive; you can use, modify, and distribute dnaio freely in commercial and private projects with minimal restrictions.
Quickstart
pip install dnaio
import dnaio
with dnaio.open("reads.fastq.gz") as f:
for record in f:
print(record.name, len(record.sequence))
Verify before relying
- Performance benchmarks comparing dnaio to other FASTQ parsers in the ecosystem
- Memory footprint for large files (e.g., multi-gigabyte FASTQ datasets)
- Interleaved paired-end parsing behavior and API details
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packagexopen |
| Maintenance | Aging 306 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 92,387 / month, #13,454 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: Science/ResearchProgramming Language :: CythonProgramming Language :: Python :: 3Topic :: Scientific/Engineering :: Bio-Informatics |
Evidence: dnaio-1.2.4-cp310-cp310-macosx_11_0_arm64.whl; dnaio-1.2.4-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp310-cp310-win_amd64.whl; dnaio-1.2.4-cp311-cp311-macosx_11_0_arm64.whl; dnaio-1.2.4-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp311-cp311-win_amd64.whl; dnaio-1.2.4-cp312-cp312-macosx_11_0_arm64.whl; dnaio-1.2.4-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp312-cp312-win_amd64.whl; dnaio-1.2.4-cp313-cp313-macosx_11_0_arm64.whl; dnaio-1.2.4-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp313-cp313-win_amd64.whl; dnaio-1.2.4-cp314-cp314-macosx_11_0_arm64.whl; dnaio-1.2.4-cp314-cp314-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp314-cp314t-macosx_11_0_arm64.whl; dnaio-1.2.4-cp314-cp314t-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; dnaio-1.2.4-cp314-cp314t-win_amd64.whl; dnaio-1.2.4-cp314-cp314-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “fastq parser python”
- dnaiodnaio reads and writes FASTQ, FASTA, and uBAM files with optimized…
- pysampysam reads, manipulates, and writes genomic data files…
- dparseParses Python dependency files (requirements.txt, conda.yml, tox.ini,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Bio-Informatics packages
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
PubChemPy is a Python wrapper around the PubChem REST API that lets you search for chemical compounds by name, substructure, or similarity, retrieve their properties, and convert between chemical file formats.
Install it if you need programmatic access to PubChem data.
See also pysam · fastar · xopen · pyfaidx · pysylph · biotite · biocommons.seqrepo · pyteomics · deepbiop · pymzml