{"categories":[{"label":"Distributed Computing","url":"https://skillfed.io/packages/category/system-distributed-computing/2"}],"enrichment":{"capability":"Dagster-spark integrates Apache Spark with Dagster's data orchestration framework, enabling you to define and run Spark-based data assets and pipelines within Dagster's declarative programming model.","skillfed_tags":["spark-integration","data-orchestration","etl"],"use_cases":["Build and orchestrate large-scale ETL pipelines using Spark within Dagster's asset framework","Define Spark transformations as reusable data assets with automatic lineage tracking","Test Spark-based computations locally during development before deploying to production","Monitor and observe Spark job execution as part of a unified Dagster data platform","Combine Spark processing with other data tools in a single orchestrated workflow"],"what_it_does":"Dagster-spark is an integration library that brings Apache Spark into Dagster's data orchestration ecosystem. It allows you to define Spark-based computations as Dagster assets\u2014Python functions that produce data assets\u2014and orchestrate them alongside other data operations within Dagster's declarative programming model. The package bridges Spark's distributed compute engine with Dagster's asset tracking, lineage, and observability features.\n\nYou use it to build data pipelines where some assets are computed via Spark jobs, while maintaining unified visibility and control through Dagster's web UI and orchestration engine. It integrates with Dagster's testing framework and supports the full development lifecycle from local testing to production deployment, letting you treat Spark workloads as first-class citizens in a broader data asset graph.","worth_installing":"Yes, if you are already using Dagster and need to integrate Spark workloads. The package is actively maintained, has low install friction, carries no known vulnerabilities, and is permissively licensed. It is worth installing as a bridge between Spark and Dagster's orchestration model, though you should verify that your target Spark version is supported."},"id":"dagster-spark","links":{"html":"https://skillfed.io/packages/dagster-spark","md":"https://skillfed.io/packages/dagster-spark.md","pypi":"https://pypi.org/project/dagster-spark/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-14","license_spdx":"Apache-2.0","license_treatment":"permissive","name":"dagster-spark","python_support":"supports_current","summary":"Package for Spark Dagster framework components."},"popularity":{"monthly_downloads":563554,"position":5979,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.29.18"}
