k-means-constrained
K-Means clustering constrained with minimum and maximum cluster size
Decision gist · record as of 2026-08-14
Yes, if you need size-constrained clustering and can tolerate higher computational cost. The package is stable (Production/Stable status), actively maintained, permissively licensed, and has no security vulnerabilities. Install friction is medium due to compiled wheels, but pre-built binaries for modern Python versions (3.10–3.13) and multiple platforms are available. Not recommended if you need vanilla k-means performance on large datasets or if cluster size constraints are not a hard requirement.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Medium install friction due to compiled wheels across multiple Python versions and platforms.
- Active maintenance with recent release (40 days ago); repository shows steady development with 236 stars and no archived status.
License · maintenance · safety
BSD 3-Clause (permissive) — BSD 3-Clause permissive license allows commercial and private use with minimal restrictions; attribution and license notice required in distributions.
last release 2026-07-05 (40 days) · last repo commit 2026-07-05 · 236 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 122,624 downloads/mo, #11,943 on PyPI
Alternatives
Verify before relying
pip install k-means-constrained
from k_means_constrained import KMeansConstrained
import numpy as np
X = np.array([[1, 2], [1, 4], [1, 0], [4, 2], [4, 4], [4, 0]])
clf = KMeansConstrained(n_clusters=2, size_min=2, size_max=5, random_state=0)
clf.fit_predict(X)- Whether performance degradation at scale is acceptable for your data size and cluster count
- Compatibility with free-threaded Python (3.14t) given ortools' GIL re-enablement behavior
What it is and what it does
k-means-constrained is a variant of k-means clustering that enforces minimum and maximum size constraints on each cluster. Instead of the standard assignment step, it formulates cluster assignment as a minimum cost flow optimization problem solved by ortools' cost-scaling push-relabel algorithm. This ensures that every cluster respects the specified size bounds while minimizing overall distance, making it useful when balanced or size-controlled partitioning is required.
The package implements a scikit-learn-compatible API and depends on ortools, scipy, numpy, six, and joblib. It trades computational speed for constraint satisfaction: the time complexity is substantially higher than vanilla k-means, scaling as O((n³c + n²c² + nc³)log(n+c)) versus O(nc) for standard k-means. The package is actively maintained, supports Python 3.10–3.13 with recent free-threading beta support, and has no known vulnerabilities.
Use it for
- Partition customer data into balanced segments for fair resource allocation or stratified analysis
- Create evenly-sized geographic clusters or facility assignments where each region must serve a minimum/maximum population
- Generate balanced train/test splits or cross-validation folds with size guarantees
- Cluster time-series or sensor data where each cluster must contain a minimum viable sample size for statistical validity
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need size-constrained clustering and can tolerate higher computational cost.
The package is stable (Production/Stable status), actively maintained, permissively licensed, and has no security vulnerabilities. Install friction is medium due to compiled wheels, but pre-built binaries for modern Python versions (3.10–3.13) and multiple platforms are available. Not recommended if you need vanilla k-means performance on large datasets or if cluster size constraints are not a hard requirement.
Install
k-means-constrained on PyPI
Before you install
Medium install friction due to compiled wheels across multiple Python versions and platforms. Active maintenance with recent release (40 days ago); repository shows steady development with 236 stars and no archived status.
License in practice
BSD 3-Clause permissive license allows commercial and private use with minimal restrictions; attribution and license notice required in distributions.
Quickstart
pip install k-means-constrained
from k_means_constrained import KMeansConstrained
import numpy as np
X = np.array([[1, 2], [1, 4], [1, 0], [4, 2], [4, 4], [4, 0]])
clf = KMeansConstrained(n_clusters=2, size_min=2, size_max=5, random_state=0)
clf.fit_predict(X)
Verify before relying
- Whether performance degradation at scale is acceptable for your data size and cluster count
- Compatibility with free-threaded Python (3.14t) given ortools' GIL re-enablement behavior
Package facts
| License | BSD 3-Clause permissive |
| Python support | Not specified |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 5 packagesortoolsscipynumpysixjoblib |
| Maintenance | Actively maintained 40 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 122,624 / month, #11,943 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseProgramming Language :: Python :: 3Programming Language :: Python :: Free Threading :: 2 - BetaTopic :: Scientific/Engineering |
Evidence: k_means_constrained-0.9.1-cp310-cp310-macosx_10_9_x86_64.whl; k_means_constrained-0.9.1-cp310-cp310-macosx_11_0_arm64.whl; k_means_constrained-0.9.1-cp310-cp310-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; k_means_constrained-0.9.1-cp310-cp310-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; k_means_constrained-0.9.1-cp310-cp310-win_amd64.whl; k_means_constrained-0.9.1-cp311-cp311-macosx_10_9_x86_64.whl; k_means_constrained-0.9.1-cp311-cp311-macosx_11_0_arm64.whl; k_means_constrained-0.9.1-cp311-cp311-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; k_means_constrained-0.9.1-cp311-cp311-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; k_means_constrained-0.9.1-cp311-cp311-win_amd64.whl; k_means_constrained-0.9.1-cp312-cp312-macosx_10_13_x86_64.whl; k_means_constrained-0.9.1-cp312-cp312-macosx_11_0_arm64.whl; k_means_constrained-0.9.1-cp312-cp312-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; k_means_constrained-0.9.1-cp312-cp312-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; k_means_constrained-0.9.1-cp312-cp312-win_amd64.whl; k_means_constrained-0.9.1-cp313-cp313-macosx_10_13_x86_64.whl; k_means_constrained-0.9.1-cp313-cp313-macosx_11_0_arm64.whl; k_means_constrained-0.9.1-cp313-cp313-manylinux2014_aarch64.manylinux_2_17_aarch64.manylinux_2_28_aarch64.whl; k_means_constrained-0.9.1-cp313-cp313-manylinux2014_x86_64.manylinux_2_17_x86_64.manylinux_2_28_x86_64.whl; k_means_constrained-0.9.1-cp313-cp313-win_amd64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “constrained k-means clustering”
- k-means-constrainedK-means clustering with enforced minimum and maximum cluster sizes,…
- kmodesImplements k-modes and k-prototypes clustering algorithms for…
- tslearntslearn provides machine learning algorithms optimized for time…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also kmodes · hdbscan · jenkspy · munkres · fastcluster · bayesian-optimization · ortools · libcuml-cu12 · cuvs-cu12