docarray
The data structure for multimodal data
Decision gist · record as of 2026-08-14
Yes. DocArray is actively maintained, has low install friction, carries a permissive Apache 2.0 license, and solves a real problem for machine learning workflows involving multimodal data. It integrates well with the Python ecosystem and is particularly valuable if you're building systems that need structured tensor and metadata handling. No known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with six runtime dependencies (pydantic, numpy, orjson, typing-inspect, types-requests, rich).
- Active maintenance with last commit on 2026-03-27 and 3125 repository stars.
License · maintenance · safety
Apache 2.0 (permissive) — Licensed under Apache 2.0 (permissive), allowing free use, modification, and distribution in both open-source and commercial projects with minimal restrictions.
last release 2025-03-21 (511 days) · last repo commit 2026-03-27 · 3,125 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 109,045 downloads/mo, #12,531 on PyPI
Alternatives
Verify before relying
pip install docarray
from docarray import BaseDoc, DocVec
from docarray.typing import ImageUrl
import numpy as np
class MyDocument(BaseDoc):
description: str
image_url: ImageUrl
tensor: np.ndarray
vec = DocVec[MyDocument]([
MyDocument(
description="A cat",
image_url="https://example.com/cat.jpg",
tensor=np.zeros((3, 224, 224)),
)
])- Whether vector database integrations (Weaviate, Qdrant, ElasticSearch, Redis, Mongo Atlas, HNSWLib) are included in base install or require optional dependencies.
- Performance characteristics when working with very large document collections or high-dimensional tensors.
- Compatibility details with specific versions of PyTorch, TensorFlow, and JAX beyond the general native support claim.
What it is and what it does
DocArray is a Python library for defining, organizing, and working with multimodal data in machine learning workflows. It provides Pydantic-based schema definitions that let you declare document types with typed fields—including tensors with explicit shapes—and then collect them into vectorized or list-based containers for batch processing. The library integrates with NumPy and other tensor frameworks, and is designed to work seamlessly with web frameworks and microservice platforms.
You use DocArray when you need to represent complex, heterogeneous data (images, text, embeddings, metadata) in a structured way that mirrors how machine learning models consume it. It handles both single documents and bulk collections: DocVec stacks tensors for efficient batch operations, while DocList preserves individual tensor structures for streaming or re-ranking. The library also supports nested document composition and can serialize data as JSON over HTTP or Protobuf over gRPC.
Use it for
- Define typed schemas for multimodal training data with tensor shape validation, then batch them for model training.
- Build API endpoints that accept and return structured multimodal documents with automatic validation.
- Organize and transmit image, text, and embedding data together in a single document structure for neural search applications.
- Compose nested document hierarchies (e.g., a document containing both image and text sub-documents) for complex data pipelines.
- Serialize multimodal collections to JSON or Protobuf for inter-service communication in microservice architectures.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
DocArray is actively maintained, has low install friction, carries a permissive Apache 2.0 license, and solves a real problem for machine learning workflows involving multimodal data. It integrates well with the Python ecosystem and is particularly valuable if you're building systems that need structured tensor and metadata handling. No known vulnerabilities.
Install
docarray on PyPI
Before you install
Low install friction with six runtime dependencies (pydantic, numpy, orjson, typing-inspect, types-requests, rich). Active maintenance with last commit on 2026-03-27 and 3125 repository stars.
License in practice
Licensed under Apache 2.0 (permissive), allowing free use, modification, and distribution in both open-source and commercial projects with minimal restrictions.
Quickstart
pip install docarray
from docarray import BaseDoc, DocVec
from docarray.typing import ImageUrl
import numpy as np
class MyDocument(BaseDoc):
description: str
image_url: ImageUrl
tensor: np.ndarray
vec = DocVec[MyDocument]([
MyDocument(
description="A cat",
image_url="https://example.com/cat.jpg",
tensor=np.zeros((3, 224, 224)),
)
])
Verify before relying
- Whether vector database integrations (Weaviate, Qdrant, ElasticSearch, Redis, Mongo Atlas, HNSWLib) are included in base install or require optional dependencies.
- Performance characteristics when working with very large document collections or high-dimensional tensors.
- Compatibility details with specific versions of PyTorch, TensorFlow, and JAX beyond the general native support claim.
Package facts
| License | Apache 2.0 permissive |
| Python support | Supports the current Python release <4.0,>=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagespydanticnumpyorjsontyping-inspecttypes-requestsrich |
| Maintenance | Actively maintained 511 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 109,045 / month, #12,531 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseLicense :: Other/Proprietary LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Unix ShellTopic :: Database :: Database Engines/ServersTopic :: Internet :: WWW/HTTP :: Indexing/SearchTopic :: Multimedia :: VideoTopic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Image RecognitionTopic :: Scientific/Engineering :: MathematicsTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: docarray-0.41.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “tensor data representation”
- docarrayDocArray provides a Python data structure for representing,…
- onnx-ironnx-ir provides an in-memory intermediate representation for ONNX…
- torchvizGenerates visual diagrams of PyTorch neural network computation…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also tensordict · tensordict-nightly · tensorly · autoray · keras · safetensors · numpydantic · xarray · torchtyping · keras-nightly