$npx skillfedfor your agent

joblibspark

Joblib Apache Spark Backend

With conditionsPyPI LibrariesReleased Apr 2025303.3K downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — joblibspark-0.6.0-py3-none-any.whl
v0.6.0 · released 2025-04-07 · 1 runtime deps: joblib

Yes, if you have a Spark cluster and need to scale joblib-parallelized training workloads. Install friction is low and the integration is straightforward. No, if you only do model inference or feature engineering in parallel—the package explicitly does not accelerate those. Verify that your PySpark version is compatible and that your use case fits the training-loop parallelization model.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires PySpark (not installed by default) and a running Spark cluster or local Spark environment; joblib>=0.14.
  • Low friction: single runtime dependency on joblib, pure Python wheel.
  • Repository is active with recent commits.

License · maintenance · safety

permissive license (permissive) — Permissive Apache license; no restrictions on commercial or proprietary use.

last release 2025-04-07 (494 days) · last repo commit 2026-03-24 · 249 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 303,304 downloads/mo, #7,814 on PyPI

Verify before relying

pip install joblibspark

from joblibspark import register_spark
from joblib import parallel_backend

register_spark()
with parallel_backend('spark', n_jobs=3):
    # joblib-parallelized code runs on Spark cluster
    pass
  • Whether the package supports Python versions beyond 3.6 and 3.7 despite classifiers listing only those.
  • Current compatibility with recent PySpark versions beyond the minimum stated.
  • Whether model inference parallelization limitations documented in the description have been addressed in recent releases.
Same gist for agents: .md · .json

What it is and what it does

Joblibspark bridges joblib's parallel execution framework and Apache Spark, allowing you to run joblib-parallelized code on a Spark cluster instead of a single machine. It works by registering Spark as a backend that joblib can dispatch tasks to, so parallel training routines and other joblib-compatible code automatically scale across cluster nodes.

The package is lightweight—it depends only on joblib—and integrates cleanly with joblib's `parallel_backend` context manager. However, it has documented limitations: it accelerates training loops but does not parallelize model inference or feature engineering steps, which continue to run locally. This makes it most useful for training workflows where the training loop itself is the bottleneck.

Use it for

  • Distribute parallel training workloads across a Spark cluster to speed up model tuning.
  • Scale joblib-based custom parallel loops to a Spark cluster without rewriting code.
  • Offload CPU-intensive training tasks from a single machine to a multi-node cluster.
  • Run joblib-parallelized code on Spark without changing application code.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you have a Spark cluster and need to scale joblib-parallelized training workloads.

Install friction is low and the integration is straightforward. No, if you only do model inference or feature engineering in parallel—the package explicitly does not accelerate those. Verify that your PySpark version is compatible and that your use case fits the training-loop parallelization model.

Install

joblibspark on PyPI

Before you install

Low friction: single runtime dependency on joblib, pure Python wheel. Repository is active with recent commits. PySpark is not bundled—you install it separately, which is typical for Spark integrations.

Requires PySpark (not installed by default) and a running Spark cluster or local Spark environment; joblib>=0.14.

License in practice

Permissive Apache license; no restrictions on commercial or proprietary use.

Quickstart

pip install joblibspark

from joblibspark import register_spark
from joblib import parallel_backend

register_spark()
with parallel_backend('spark', n_jobs=3):
    # joblib-parallelized code runs on Spark cluster
    pass

Verify before relying

  • Whether the package supports Python versions beyond 3.6 and 3.7 despite classifiers listing only those.
  • Current compatibility with recent PySpark versions beyond the minimum stated.
  • Whether model inference parallelization limitations documented in the description have been addressed in recent releases.

Package facts

Licensepermissive license permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
joblib
MaintenanceActively maintained 494 days since the last release
Last repo commit
First released
Downloads303,304 / month, #7,814 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Topic :: Scientific/EngineeringTopic :: Software Development :: LibrariesTopic :: Utilities

Evidence: joblibspark-0.6.0-py3-none-any.whl

Tags

Capabilities
spark backend for joblibdistributed parallelization sparkparallel joblib sparkspark cluster parallelizationjoblib spark integration
Topics
spark-integrationdistributed-computing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “spark backend for joblib”

  • joblibsparkRegisters Apache Spark as a distributed backend for joblib's parallel…
  • spark-sklearnDistributes scikit-learn model training and hyperparameter search…
  • mrmr-selectionImplements the mRMR (minimum Redundancy - Maximum Relevance) feature…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also spark-sklearn · pyspark-client · pyspark · livy · tqdm-joblib · gamma-pytools · pyspark-pandas · snowpark-connect · koalas · pyspark-extension