synapseml
Synapse Machine Learning
Decision gist · record as of 2026-08-14
Yes, if you are working with Spark and need to build distributed ML pipelines. SynapseML is actively maintained, has no known vulnerabilities, and offers a rich set of ML capabilities that integrate seamlessly with existing Spark workflows. The main condition is that you must already have Spark 3.4+, Scala 2.12, and Python 3.8+ configured; if you lack a Spark environment, the setup overhead may be substantial. For teams already using Spark, it is a natural choice for scaling ML work.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Apache Spark 3.4+, Scala 2.12, and Python 3.8+ to be installed and configured in your environment before using SynapseML.
- Installation is straightforward (pure Python wheel), and the project is actively maintained with recent commits.
- However, it requires Spark 3.4+, Scala 2.12, and Python 3.8+ as runtime dependencies outside the PyPI package itself, which may add setup complexity depending on your environment.
License · maintenance · safety
MIT (permissive) — Licensed under MIT (permissive), so you can use, modify, and distribute SynapseML with minimal legal restrictions in commercial and open-source projects.
last release 2026-04-07 (129 days) · last repo commit 2026-08-14 · 5,236 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,800,776 downloads/mo, #3,543 on PyPI
Alternatives
Verify before relying
pip install synapseml
# SynapseML integrates with Apache Spark for distributed ML pipelines
# See documentation for usage patterns with your Spark environment- Specific performance benchmarks or scalability limits (e.g., maximum cluster size tested, throughput on typical workloads).
- Compatibility matrix with specific Spark versions beyond the stated 3.4+ requirement.
- Whether all advertised features (Vowpal Wabbit, LightGBM, ONNX, HTTP on Spark) are equally mature or if some are experimental.
- Concrete code examples showing how to import and use SynapseML modules in practice.
What it is and what it does
SynapseML is a distributed machine learning library that extends Apache Spark with high-level APIs for building scalable ML pipelines. It abstracts over text analytics, computer vision, anomaly detection, deep learning, and other ML tasks, allowing you to compose these capabilities into workflows that run on single-node, multi-node, or elastically resizable clusters. The library shares the same API conventions as SparkML/MLLib, so it integrates naturally into existing Spark workflows.
The package is designed to handle data wherever it lives—across different databases, file systems, and cloud stores—and is usable from Python, R, Scala, Java, and .NET. It includes specialized modules for sparse text analytics, gradient boosting, ONNX model serving, and HTTP-based model deployment. The main constraint is that you must have Spark 3.4+, Scala 2.12, and Python 3.8+ already installed and configured.
Use it for
- Train and deploy text analytics models on multi-terabyte datasets using distributed algorithms.
- Build computer vision pipelines that process images at scale across a distributed cluster.
- Detect anomalies in large time-series or streaming data using distributed ML algorithms.
- Integrate cognitive services into Spark ML workflows for big-data AI applications.
- Serve trained Spark models as low-latency web services.
- Train gradient boosted decision trees on large datasets using distributed computing.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are working with Spark and need to build distributed ML pipelines.
SynapseML is actively maintained, has no known vulnerabilities, and offers a rich set of ML capabilities that integrate seamlessly with existing Spark workflows. The main condition is that you must already have Spark 3.4+, Scala 2.12, and Python 3.8+ configured; if you lack a Spark environment, the setup overhead may be substantial. For teams already using Spark, it is a natural choice for scaling ML work.
Install
synapseml on PyPI
Before you install
Installation is straightforward (pure Python wheel), and the project is actively maintained with recent commits. However, it requires Spark 3.4+, Scala 2.12, and Python 3.8+ as runtime dependencies outside the PyPI package itself, which may add setup complexity depending on your environment.
Requires Apache Spark 3.4+, Scala 2.12, and Python 3.8+ to be installed and configured in your environment before using SynapseML.
License in practice
Licensed under MIT (permissive), so you can use, modify, and distribute SynapseML with minimal legal restrictions in commercial and open-source projects.
Quickstart
pip install synapseml
# SynapseML integrates with Apache Spark for distributed ML pipelines
# See documentation for usage patterns with your Spark environment
Verify before relying
- Specific performance benchmarks or scalability limits (e.g., maximum cluster size tested, throughput on typical workloads).
- Compatibility matrix with specific Spark versions beyond the stated 3.4+ requirement.
- Whether all advertised features (Vowpal Wabbit, LightGBM, ONNX, HTTP on Spark) are equally mature or if some are experimental.
- Concrete code examples showing how to import and use SynapseML modules in practice.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 129 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,800,776 / month, #3,543 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 2Programming Language :: Python :: 3Topic :: Software Development :: Libraries |
Evidence: synapseml-1.1.3-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “scalable ml pipelines”
- synapsemlSynapseML is a distributed machine learning library built on Apache…
- dask-mlDask-ML provides distributed and parallel machine learning by…
- kfpKubeflow Pipelines is a Python SDK for defining, deploying, and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also pyspark · azure-synapse-spark · dask-ml · spark-sklearn · mleap · spark-nlp · pyspark-client · azureml-sdk · clearml · pyspark-pandas