$npx skillfedfor your agent

spark-nlp

John Snow Labs Spark NLP is a natural language processing library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines, that scale easily in a distributed environment.

With conditionsPyPI Scientific/EngineeringReleased Jun 20261.2M downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — spark_nlp-6.4.2-py2.py3-none-any.whl
v6.4.2 · released 2026-06-24

Yes, if you need production-grade NLP at scale on Spark. The library is actively maintained with no security vulnerabilities and offers a comprehensive suite of pretrained models and tasks. Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Java 8 or 11 (Oracle or OpenJDK) and Apache Spark 3.0+ to be installed and configured in your environment.
  • Installation is straightforward (low friction) with no runtime dependencies to manage.
  • The package is actively maintained with a recent release 51 days ago.

License · maintenance · safety

permissive license (permissive) — Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.

last release 2026-06-24 (51 days) · last repo commit 2026-08-13 · 4,155 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,180,010 downloads/mo, #4,260 on PyPI

Verify before relying

pip install spark-nlp==6.4.2

from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline

pipeline = PretrainedPipeline('explain_document_dl', lang='en')
result = pipeline.annotate("The Mona Lisa is a 16th century oil painting created by Leonardo.")
  • Whether pretrained model downloads are automatic or require separate setup steps
  • Memory and disk requirements for downloading and caching pretrained models
  • Performance characteristics on single-machine vs. true distributed clusters
  • Exact Python version support (classifiers list 3.6–3.9 but requires_python is unspecified)
Same gist for agents: .md · .json

What it is and what it does

Spark NLP is a natural language processing library built on top of Apache Spark that brings NLP and machine learning to distributed environments. It provides access to pretrained pipelines and models across multiple languages, supporting tasks like tokenization, part-of-speech tagging, named entity recognition, sentiment analysis, machine translation, question answering, and text generation. The library integrates state-of-the-art transformer models and can import models from TensorFlow, ONNX, and OpenVINO frameworks.

The package is designed for production use and scales across distributed Spark clusters. It supports multiple programming languages through the JVM ecosystem and offers specialized variants for GPU acceleration, Apple Silicon, and AArch64 architectures. The library requires Java 8 or 11 and Apache Spark 3.0 or later.

Use it for

  • Build production NLP pipelines that scale across Spark clusters for large-scale text processing
  • Extract named entities, perform sentiment analysis, or classify documents using pretrained models
  • Translate text between languages or generate summaries and answers from documents
  • Integrate transformer models into Spark ML workflows for feature engineering
  • Process multilingual text data in a single distributed pipeline

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need production-grade NLP at scale on Spark.

The library is actively maintained with no security vulnerabilities and offers a comprehensive suite of pretrained models and tasks. Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.

Install

spark-nlp on PyPI

Before you install

Installation is straightforward (low friction) with no runtime dependencies to manage. The package is actively maintained with a recent release 51 days ago. Requires Java 8 or 11 and Apache Spark 3.0+ to be present in your environment.

Requires Java 8 or 11 (Oracle or OpenJDK) and Apache Spark 3.0+ to be installed and configured in your environment.

License in practice

Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.

Quickstart

pip install spark-nlp==6.4.2

from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline

pipeline = PretrainedPipeline('explain_document_dl', lang='en')
result = pipeline.annotate("The Mona Lisa is a 16th century oil painting created by Leonardo.")

Verify before relying

  • Whether pretrained model downloads are automatic or require separate setup steps
  • Memory and disk requirements for downloading and caching pretrained models
  • Performance characteristics on single-machine vs. true distributed clusters
  • Exact Python version support (classifiers list 3.6–3.9 but requires_python is unspecified)

Package facts

Licensepermissive license permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 51 days since the last release
Last repo commit
First released
Downloads1,180,010 / month, #4,260 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: OS IndependentOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: Build ToolsTopic :: Software Development :: InternationalizationTopic :: Software Development :: Libraries :: Python ModulesTopic :: Software Development :: LocalizationTopic :: Text Processing :: LinguisticTyping :: Typed

Evidence: spark_nlp-6.4.2-py2.py3-none-any.whl

Tags

Capabilities
distributed NLP on Sparkpretrained NLP pipelinesnamed entity recognitionsentiment analysis at scalemachine translationtext embeddingstransformer models for NLP
Topics
distributed-nlpspark-mltransformers
PyPI keywords
NLPsparkvisionspeechdeeplearningtransformertensorflowBERTGPT-2Wav2Vec2ViT

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “distributed NLP on Spark”

  • spark-nlpSpark NLP provides distributed natural language processing on Apache…
  • koalasKoalas implements the pandas DataFrame API on top of Apache Spark,…
  • pyspark-pandasProvides tools for distributing Pandas DataFrames and Series across…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also flair · torchtext · spacy · textblob · keras-nlp · polyglot · google-cloud-language · urduhack · mleap · synapseml