{"categories":[{"label":"Build Tools","url":"https://skillfed.io/packages/category/software-development-build-tools/6"}],"enrichment":{"capability":"PySparkling provides Python bindings to train and score H2O-3 and Sparkling Water machine learning models on Apache Spark, with support for MOJO model formats.","skillfed_tags":["spark-integration","distributed-ml","mojo-models"],"use_cases":["Train H2O machine learning models on large distributed datasets using Spark's cluster computing.","Score pre-trained MOJO models in batch or streaming Spark jobs for inference at scale.","Integrate Driverless AI MOJO models into Spark-based data pipelines for automated predictions.","Build end-to-end machine learning workflows combining Spark data preparation with H2O model training.","Deploy H2O models to production Spark clusters for real-time or batch scoring on big data."],"what_it_does":"PySparkling is a Python library that bridges H2O's machine learning engine with Apache Spark, enabling distributed model training and scoring across Spark clusters. It provides APIs to work with H2O-3 models and MOJO (Model Object, Optimized) artifacts, which are H2O's portable, language-agnostic model format. The package also supports scoring Driverless AI MOJO models.\n\nThe library is designed for data scientists and engineers working in big-data environments who need to integrate H2O's statistical and machine learning capabilities into Spark-based workflows. It has no runtime dependencies listed in the fact sheet, meaning it relies entirely on an external Spark installation and H2O runtime to function. The package is classified as Production/Stable and supports Python 3.6 through 3.10, though the last release was in November 2024.","worth_installing":"Yes, with conditions. PySparkling is worth installing if you are already committed to an Apache Spark infrastructure and need H2O's machine learning capabilities in that environment. The permissive Apache license and Production/Stable status support production use. However, high install friction, aging maintenance (633 days since last release), and the requirement for a pre-configured Spark cluster mean this is not a lightweight addition\u2014evaluate whether H2O's specific algorithms justify the operational overhead in your stack."},"id":"h2o-pysparkling-3-1","links":{"html":"https://skillfed.io/packages/h2o-pysparkling-3-1","md":"https://skillfed.io/packages/h2o-pysparkling-3-1.md","pypi":"https://pypi.org/project/h2o-pysparkling-3-1/"},"maintenance":{"status":"aging"},"meta":{"latest_release":"2024-11-19","license_spdx":null,"license_treatment":"permissive","name":"h2o-pysparkling-3.1","python_support":"unspecified","summary":"Sparkling Water integrates H2O's Fast Scalable Machine Learning with Spark"},"popularity":{"monthly_downloads":79470,"position":14354,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"3.46.0.6.post1"}
