{"categories":[{"label":"Libraries","url":"https://skillfed.io/packages/category/software-development-libraries/3"}],"enrichment":{"capability":"SynapseML is a distributed machine learning library built on Apache Spark that provides APIs for text analytics, computer vision, anomaly detection, and other ML tasks, designed to scale across single-node and multi-node clusters.","skillfed_tags":["distributed-ml","spark-integration","big-data"],"use_cases":["Train and deploy text analytics models on multi-terabyte datasets using distributed algorithms.","Build computer vision pipelines that process images at scale across a distributed cluster.","Detect anomalies in large time-series or streaming data using distributed ML algorithms.","Integrate cognitive services into Spark ML workflows for big-data AI applications.","Serve trained Spark models as low-latency web services.","Train gradient boosted decision trees on large datasets using distributed computing."],"what_it_does":"SynapseML is a distributed machine learning library that extends Apache Spark with high-level APIs for building scalable ML pipelines. It abstracts over text analytics, computer vision, anomaly detection, deep learning, and other ML tasks, allowing you to compose these capabilities into workflows that run on single-node, multi-node, or elastically resizable clusters. The library shares the same API conventions as SparkML/MLLib, so it integrates naturally into existing Spark workflows.\n\nThe package is designed to handle data wherever it lives\u2014across different databases, file systems, and cloud stores\u2014and is usable from Python, R, Scala, Java, and .NET. It includes specialized modules for sparse text analytics, gradient boosting, ONNX model serving, and HTTP-based model deployment. The main constraint is that you must have Spark 3.4+, Scala 2.12, and Python 3.8+ already installed and configured.","worth_installing":"Yes, if you are working with Spark and need to build distributed ML pipelines. SynapseML is actively maintained, has no known vulnerabilities, and offers a rich set of ML capabilities that integrate seamlessly with existing Spark workflows. The main condition is that you must already have Spark 3.4+, Scala 2.12, and Python 3.8+ configured; if you lack a Spark environment, the setup overhead may be substantial. For teams already using Spark, it is a natural choice for scaling ML work."},"id":"synapseml","links":{"html":"https://skillfed.io/packages/synapseml","md":"https://skillfed.io/packages/synapseml.md","pypi":"https://pypi.org/project/synapseml/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-04-07","license_spdx":null,"license_treatment":"permissive","name":"synapseml","python_support":"unspecified","summary":"Synapse Machine Learning"},"popularity":{"monthly_downloads":1800776,"position":3543,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"1.1.3"}
