tiledbsoma
Python API for efficient storage and retrieval of single-cell data using TileDB
Decision gist · record as of 2026-08-14
Yes, if you work with single-cell genomic data and need standardized, efficient storage. The package is actively maintained, supports modern Python versions (3.9–3.13), has no known vulnerabilities, and integrates well with the established scanpy and anndata ecosystem. Medium install friction is manageable for most development environments. Not necessary if you are already satisfied with your current single-cell data storage and don't require SOMA compliance or TileDB's specific performance characteristics.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later.
- On older CPU architectures without AVX2 support, source installation from the repository may be necessary instead of using pre-compiled wheels.
- Medium install friction due to compiled binary wheels for multiple Python versions and architectures.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for both academic and production environments.
last release 2026-01-27 (199 days) · last repo commit 2026-07-28 · 131 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 100,806 downloads/mo, #12,973 on PyPI
Alternatives
Verify before relying
pip install tiledbsoma
import tiledbsoma as soma
# Access SOMA data via TileDB backend with platform_config for TileDB-specific settings- Whether the package requires a TileDB server or library installation separate from the Python package itself
- Performance characteristics and scalability limits for typical single-cell datasets
- Compatibility with specific versions of scanpy, anndata, or other bioinformatics tools in the dependency chain
What it is and what it does
TileDB-SOMA provides a Python API for the Unified Single-cell Data Model (SOMA), an open standard for organizing and accessing single-cell genomic data. It uses TileDB, a columnar array storage engine, as its backend to enable efficient storage and retrieval of large single-cell datasets. The package implements the full SOMA specification, allowing researchers and bioinformaticians to work with standardized single-cell data formats across different tools and platforms.
The package integrates with the popular single-cell Python ecosystem—it depends on anndata, scanpy, pandas, numpy, and pyarrow—making it a bridge between TileDB's storage capabilities and existing bioinformatics workflows. It supports platform-specific configuration through a TypeScript-style interface, allowing fine-grained control over TileDB storage parameters like filters, cell order, and capacity settings.
Use it for
- Store and query large single-cell RNA-seq datasets in a standardized, interoperable format using TileDB's efficient columnar storage.
- Integrate single-cell data workflows with tools like scanpy and anndata while maintaining SOMA compliance across different analysis pipelines.
- Configure TileDB-specific storage options (filters, dimensions, compression) for optimized performance on custom single-cell datasets.
- Access single-cell data through a standardized API that abstracts away storage backend details, enabling portability between implementations.
- Build reproducible bioinformatics pipelines that rely on a unified data model for single-cell genomics across multiple research groups.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with single-cell genomic data and need standardized, efficient storage.
The package is actively maintained, supports modern Python versions (3.9–3.13), has no known vulnerabilities, and integrates well with the established scanpy and anndata ecosystem. Medium install friction is manageable for most development environments. Not necessary if you are already satisfied with your current single-cell data storage and don't require SOMA compliance or TileDB's specific performance characteristics.
Install
tiledbsoma on PyPI
Before you install
Medium install friction due to compiled binary wheels for multiple Python versions and architectures. Active maintenance with recent releases. Requires 10 runtime dependencies including numpy, pandas, pyarrow, and scanpy. Pre-built wheels available for macOS and Linux; source installation may be needed on older processors lacking AVX2 support.
Requires Python 3.9 or later. On older CPU architectures without AVX2 support, source installation from the repository may be necessary instead of using pre-compiled wheels.
License in practice
MIT license permits commercial and private use with minimal restrictions, making it suitable for both academic and production environments.
Quickstart
pip install tiledbsoma
import tiledbsoma as soma
# Access SOMA data via TileDB backend with platform_config for TileDB-specific settings
Verify before relying
- Whether the package requires a TileDB server or library installation separate from the Python package itself
- Performance characteristics and scalability limits for typical single-cell datasets
- Compatibility with specific versions of scanpy, anndata, or other bioinformatics tools in the dependency chain
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 10 packagesanndataattrsmore-itertoolsnumpypandaspyarrowscanpyscipysomacoretyping-extensions |
| Maintenance | Actively maintained 199 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 100,806 / month, #12,973 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxOperating System :: UnixProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Bio-Informatics |
Evidence: tiledbsoma-2.3.0-cp310-cp310-macosx_13_0_arm64.whl; tiledbsoma-2.3.0-cp310-cp310-macosx_13_0_x86_64.whl; tiledbsoma-2.3.0-cp310-cp310-manylinux_2_28_aarch64.whl; tiledbsoma-2.3.0-cp310-cp310-manylinux_2_28_x86_64.whl; tiledbsoma-2.3.0-cp311-cp311-macosx_13_0_arm64.whl; tiledbsoma-2.3.0-cp311-cp311-macosx_13_0_x86_64.whl; tiledbsoma-2.3.0-cp311-cp311-manylinux_2_28_aarch64.whl; tiledbsoma-2.3.0-cp311-cp311-manylinux_2_28_x86_64.whl; tiledbsoma-2.3.0-cp312-cp312-macosx_13_0_arm64.whl; tiledbsoma-2.3.0-cp312-cp312-macosx_13_0_x86_64.whl; tiledbsoma-2.3.0-cp312-cp312-manylinux_2_28_aarch64.whl; tiledbsoma-2.3.0-cp312-cp312-manylinux_2_28_x86_64.whl; tiledbsoma-2.3.0-cp313-cp313-macosx_13_0_arm64.whl; tiledbsoma-2.3.0-cp313-cp313-macosx_13_0_x86_64.whl; tiledbsoma-2.3.0-cp313-cp313-manylinux_2_28_aarch64.whl; tiledbsoma-2.3.0-cp313-cp313-manylinux_2_28_x86_64.whl; tiledbsoma-2.3.0-cp39-cp39-macosx_13_0_arm64.whl; tiledbsoma-2.3.0-cp39-cp39-macosx_13_0_x86_64.whl; tiledbsoma-2.3.0-cp39-cp39-manylinux_2_28_aarch64.whl; tiledbsoma-2.3.0-cp39-cp39-manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “single-cell data storage”
- tiledbsomaTileDB-SOMA is a Python implementation of the SOMA API specification…
- somacoreProvides the Python reference implementation of the SOMA…
- anndataanndata handles annotated data matrices in memory and on disk with…
Give your agent the search over MCP, or paste the wish link into any chat.
More Bio-Informatics packages
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Biopython provides Python tools for computational molecular biology, including sequence analysis, structure parsing, database access, and phylogenetic tree manipulation.
However, verify that the custom Biopython License Agreement aligns with your project's licensing requirements before committing to it in production or proprietary work.
Client library for the Firecrawl API that scrapes, crawls, and searches the web, returning clean Markdown or structured data; also indexes research papers from PubMed, bioRxiv, medRxiv, and arXiv.
A self-balancing interval tree data structure that stores and queries overlapping or enveloped ranges, supporting point lookups, range overlaps, and range envelopment queries.
Install it if you need to store and query overlapping or enveloped ranges; the self-balancing design and rich query interface make it significantly easier than…
Albumentations applies image transformations to training data, supporting classification, segmentation, object detection, and pose estimation with a unified API for images, masks, bounding boxes, and keypoints.
Install it if you need a unified, production-grade augmentation API for computer vision tasks.
PubChemPy is a Python wrapper around the PubChem REST API that lets you search for chemical compounds by name, substructure, or similarity, retrieve their properties, and convert between chemical file formats.
Install it if you need programmatic access to PubChem data.
See also cellxgene-census · somacore · tiledb · mudata · scvi-tools · scanpy · anndata · pyranges · latch · biocommons.seqrepo