dask-expr
High Level Expressions for Dask
Decision gist · record as of 2026-08-14
Yes. Dask Expressions is the default backend for dask.DataFrame as of version 2024.3.0, making it the standard choice for new Dask DataFrame work. It offers query optimization with low install friction and permissive licensing. Maintenance is dormant but the package is stable and widely used. Install it if you are using Dask DataFrames; it is the recommended path forward.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low install friction with a single wheel distribution.
- Maintenance is dormant (570 days since last release), though the repository remains active with a recent commit on 2025-01-21.
License · maintenance · safety
BSD (permissive) — BSD license is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most projects.
last release 2025-01-21 (570 days) · last repo commit 2025-01-21 · 89 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 10,958,153 downloads/mo, #1,422 on PyPI
Alternatives
Verify before relying
pip install dask-expr
import dask_expr as dx
df = dx.datasets.timeseries()
result = df.groupby("name").x.mean().compute()- Whether named GroupBy Aggregations (noted as missing) are critical for your workflow.
- Performance gains over standard Dask DataFrame for your specific query patterns.
What it is and what it does
Dask Expressions is a rewrite of Dask DataFrame that adds query optimization by representing user operations as an expression tree before execution. Instead of executing operations immediately, the library builds a tree structure encoding the computation, optimizes it (e.g., fusing operations, reordering steps), and then executes the optimized plan. It is designed as a drop-in replacement and has been the default backend for dask.DataFrame since version 2024.3.0.
The package depends only on dask and installs with low friction. It covers nearly the full Dask DataFrame API, with the exception of named GroupBy Aggregations. It supports modern Python versions (3.10 through 3.13) and is intended for developers and researchers working with distributed data processing at scale.
Use it for
- Optimize large distributed dataframe queries by leveraging automatic expression tree optimization before execution.
- Migrate from standard Dask DataFrame to a more efficient backend without changing existing code.
- Process time-series and grouped data with better performance through query planning.
- Build data pipelines where query optimization can reduce computation time and memory usage.
- Work with pandas-like APIs on distributed datasets while benefiting from automatic optimization.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Dask Expressions is the default backend for dask.DataFrame as of version 2024.3.0, making it the standard choice for new Dask DataFrame work. It offers query optimization with low install friction and permissive licensing. Maintenance is dormant but the package is stable and widely used. Install it if you are using Dask DataFrames; it is the recommended path forward.
Install
dask-expr on PyPI
Before you install
Low install friction with a single wheel distribution. Maintenance is dormant (570 days since last release), though the repository remains active with a recent commit on 2025-01-21. This is the default backend for dask.DataFrame since version 2024.3.0, so stability is established.
Requires Python 3.10 or later.
License in practice
BSD license is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most projects.
Quickstart
pip install dask-expr
import dask_expr as dx
df = dx.datasets.timeseries()
result = df.groupby("name").x.mean().compute()
Verify before relying
- Whether named GroupBy Aggregations (noted as missing) are critical for your workflow.
- Performance gains over standard Dask DataFrame for your specific query patterns.
Package facts
| License | BSD permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagedask |
| Maintenance | Dormant 570 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 10,958,153 / month, #1,422 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: System :: Distributed Computing |
Evidence: dask_expr-2.0.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “dask dataframe query optimization”
- dask-exprDask Expressions provides query optimization for Dask DataFrames by…
- swifterSwifter applies functions to DataFrames and Series using automatic…
- dask-cudf-cu12Dask cuDF extends Dask DataFrame with a GPU-accelerated backend,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also qpd · flox · swifter · dask-cudf-cu12 · polars-runtime-compat · polars · dask-geopandas · accumulation-tree · substrait · etuples