distributed
Distributed scheduler for Dask
Decision gist · record as of 2026-08-14
Yes. Distributed is production-stable, actively maintained, and widely used for scaling Dask workloads. Install friction is low, dependencies are well-established, and there are no known vulnerabilities. Choose it if you need to parallelize or distribute computation beyond a single machine; if you only need local parallelism, Dask alone may suffice.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Typically used with Dask; standalone use is possible but limited.
- Low install friction with a pure-Python wheel.
License · maintenance · safety
BSD-3-Clause (permissive) — BSD-3-Clause is permissive; you may use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.
last release 2026-07-14 (31 days) · last repo commit 2026-08-14 · 1,676 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 8,413,098 downloads/mo, #1,625 on PyPI
Alternatives
Verify before relying
pip install distributed
from distributed import Client
client = Client()
result = client.submit(lambda x: x + 1, 1).result()- Whether the package works reliably on all supported Python versions (3.10–3.14) in production environments.
- Performance characteristics and scalability limits for different cluster sizes and workload types.
What it is and what it does
Distributed is a scheduler and runtime that coordinates parallel and distributed computation for Dask. It manages task scheduling, data movement, and fault tolerance across a cluster of workers, allowing you to run computations on multiple machines or cores as if they were a single pool. The package handles the low-level coordination—worker lifecycle, task graph execution, communication—so you focus on expressing your computation in Dask.
You typically use it by creating a Client that connects to a cluster (local or remote), then submitting Dask graphs or collections through that client. It depends on Tornado for async networking, Cloudpickle for serialization, and several utility libraries (toolz, sortedcontainers, locket, msgpack) to manage scheduling state and inter-process communication. The package is production-stable and actively maintained.
Use it for
- Scale data processing pipelines across a cluster without rewriting code for distributed execution.
- Run machine learning training or inference jobs in parallel across multiple GPUs or nodes.
- Execute long-running batch computations that would timeout or exhaust memory on a single machine.
- Coordinate complex workflows with dependencies between tasks across a heterogeneous cluster.
- Monitor and debug distributed workloads through the scheduler's built-in diagnostics and web dashboard.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Distributed is production-stable, actively maintained, and widely used for scaling Dask workloads. Install friction is low, dependencies are well-established, and there are no known vulnerabilities. Choose it if you need to parallelize or distribute computation beyond a single machine; if you only need local parallelism, Dask alone may suffice.
Install
distributed on PyPI
Before you install
Low install friction with a pure-Python wheel. Active maintenance with a recent release (31 days ago) and ongoing commits. Supports current Python versions (3.10–3.14).
Requires Python 3.10 or later. Typically used with Dask; standalone use is possible but limited.
License in practice
BSD-3-Clause is permissive; you may use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.
Quickstart
pip install distributed
from distributed import Client
client = Client()
result = client.submit(lambda x: x + 1, 1).result()
Verify before relying
- Whether the package works reliably on all supported Python versions (3.10–3.14) in production environments.
- Performance characteristics and scalability limits for different cluster sizes and workload types.
Package facts
| License | BSD-3-Clause permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 14 packagesclickcloudpickledaskjinja2locketmsgpackpackagingpsutilpyyamlsortedcontainerstblibtoolztornadozict |
| Maintenance | Actively maintained 31 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 8,413,098 / month, #1,625 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Scientific/EngineeringTopic :: System :: Distributed Computing |
Evidence: distributed-2026.7.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “parallel computation framework”
- distributedDistributed provides a scheduler and runtime for parallel and…
- drjitDr.Jit is a just-in-time compiler for differentiable and ordinary…
- openmpiProvides a Python interface to Open MPI, a high-performance…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also coiled · dask · dask-jobqueue · mitogen · anyscale · prefect-dask · dask-image · lithops · ipyparallel · dask-geopandas