flox
GroupBy operations for dask.array
Decision gist · record as of 2026-08-14
Yes. Flox is actively maintained, has no security vulnerabilities, and low install friction. It directly addresses performance needs in GroupBy operations for scientific computing. The Apache-2.0 license is permissive. Install it if you need efficient grouped reductions on arrays; the Beta status is not a barrier given the active maintenance and NASA-funded backing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.11 or later.
- Low install friction with a pure Python wheel and six well-established dependencies.
- Active maintenance with recent commits; marked Beta but supported on current Python versions (3.11–3.14).
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows use in most projects without restriction, including commercial applications.
last release 2026-02-26 (169 days) · last repo commit 2026-07-29 · 137 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 352,795 downloads/mo, #7,309 on PyPI
Alternatives
Verify before relying
pip install flox
import flox
# GroupBy reduction using groupby_reduce
result = flox.groupby_reduce(array, by_array, "mean")- Whether custom Aggregation support is fully functional (description notes it 'might not be fully functional at the moment').
- Performance benchmarks or typical speedup factors compared to alternative GroupBy approaches.
- Whether xarray integration has progressed beyond the ongoing work mentioned in the description.
What it is and what it does
Flox is a library for performing GroupBy reduction operations efficiently on arrays with dask and xarray. It was created to address performance bottlenecks in GroupBy operations and to bring fast aggregations to the dask ecosystem. The package provides two main interfaces: a pure array function (groupby_reduce) and an xarray-specific function (xarray_reduce), both supporting common reductions like mean, sum, and count, as well as custom aggregations defined via an Aggregation class.
The library depends on pandas, numpy, scipy, packaging, toolz, and numpy_groupies to deliver its functionality. It is actively maintained, supports Python 3.11–3.14, and carries an Apache-2.0 license. With low install friction and no known security vulnerabilities, it is suitable for scientific computing workflows that require scalable GroupBy operations on large, chunked arrays.
Use it for
- Compute grouped statistics on large arrays without materializing the full dataset in memory.
- Perform GroupBy reductions on multi-dimensional scientific data in a distributed setting.
- Define custom aggregation logic for domain-specific GroupBy operations on chunked arrays.
- Integrate fast GroupBy operations into Earth science data analysis pipelines.
- Reduce intermediate results from blockwise operations using combine and finalize strategies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Flox is actively maintained, has no security vulnerabilities, and low install friction. It directly addresses performance needs in GroupBy operations for scientific computing. The Apache-2.0 license is permissive. Install it if you need efficient grouped reductions on arrays; the Beta status is not a barrier given the active maintenance and NASA-funded backing.
Install
flox on PyPI
Before you install
Low install friction with a pure Python wheel and six well-established dependencies. Active maintenance with recent commits; marked Beta but supported on current Python versions (3.11–3.14).
Requires Python 3.11 or later.
License in practice
Apache-2.0 permissive license allows use in most projects without restriction, including commercial applications.
Quickstart
pip install flox
import flox
# GroupBy reduction using groupby_reduce
result = flox.groupby_reduce(array, by_array, "mean")
Verify before relying
- Whether custom Aggregation support is fully functional (description notes it 'might not be fully functional at the moment').
- Performance benchmarks or typical speedup factors compared to alternative GroupBy approaches.
- Whether xarray integration has progressed beyond the ongoing work mentioned in the description.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagespandaspackagingnumpynumpy_groupiestoolzscipy |
| Maintenance | Actively maintained 169 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 352,795 / month, #7,309 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: flox-0.11.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “groupby reductions”
- floxFlox provides fast GroupBy reduction operations for arrays using dask…
- numpy-groupiesProvides optimized group-indexing operations on arrays, with…
- numbaggNumbagg provides fast N-dimensional aggregation and moving-window…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also numpy-groupies · window-ops · dask-expr · xarray · accumulation-tree · rasterix · dask · odc-loader · linopy · qpd