spark-nlp
John Snow Labs Spark NLP is a natural language processing library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines, that scale easily in a distributed environment.
Decision gist · record as of 2026-08-14
Yes, if you need production-grade NLP at scale on Spark. The library is actively maintained with no security vulnerabilities and offers a comprehensive suite of pretrained models and tasks. Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Java 8 or 11 (Oracle or OpenJDK) and Apache Spark 3.0+ to be installed and configured in your environment.
- Installation is straightforward (low friction) with no runtime dependencies to manage.
- The package is actively maintained with a recent release 51 days ago.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.
last release 2026-06-24 (51 days) · last repo commit 2026-08-13 · 4,155 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,180,010 downloads/mo, #4,260 on PyPI
Alternatives
Verify before relying
pip install spark-nlp==6.4.2
from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline
pipeline = PretrainedPipeline('explain_document_dl', lang='en')
result = pipeline.annotate("The Mona Lisa is a 16th century oil painting created by Leonardo.")- Whether pretrained model downloads are automatic or require separate setup steps
- Memory and disk requirements for downloading and caching pretrained models
- Performance characteristics on single-machine vs. true distributed clusters
- Exact Python version support (classifiers list 3.6–3.9 but requires_python is unspecified)
What it is and what it does
Spark NLP is a natural language processing library built on top of Apache Spark that brings NLP and machine learning to distributed environments. It provides access to pretrained pipelines and models across multiple languages, supporting tasks like tokenization, part-of-speech tagging, named entity recognition, sentiment analysis, machine translation, question answering, and text generation. The library integrates state-of-the-art transformer models and can import models from TensorFlow, ONNX, and OpenVINO frameworks.
The package is designed for production use and scales across distributed Spark clusters. It supports multiple programming languages through the JVM ecosystem and offers specialized variants for GPU acceleration, Apple Silicon, and AArch64 architectures. The library requires Java 8 or 11 and Apache Spark 3.0 or later.
Use it for
- Build production NLP pipelines that scale across Spark clusters for large-scale text processing
- Extract named entities, perform sentiment analysis, or classify documents using pretrained models
- Translate text between languages or generate summaries and answers from documents
- Integrate transformer models into Spark ML workflows for feature engineering
- Process multilingual text data in a single distributed pipeline
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need production-grade NLP at scale on Spark.
The library is actively maintained with no security vulnerabilities and offers a comprehensive suite of pretrained models and tasks. Install only if you already have Apache Spark 3.0+ and Java 8 or 11 in your environment; it is not suitable for lightweight single-machine NLP work.
Install
spark-nlp on PyPI
Before you install
Installation is straightforward (low friction) with no runtime dependencies to manage. The package is actively maintained with a recent release 51 days ago. Requires Java 8 or 11 and Apache Spark 3.0+ to be present in your environment.
Requires Java 8 or 11 (Oracle or OpenJDK) and Apache Spark 3.0+ to be installed and configured in your environment.
License in practice
Licensed under Apache Software License (permissive), allowing commercial and private use with minimal restrictions.
Quickstart
pip install spark-nlp==6.4.2
from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline
pipeline = PretrainedPipeline('explain_document_dl', lang='en')
result = pipeline.annotate("The Mona Lisa is a 16th century oil painting created by Leonardo.")
Verify before relying
- Whether pretrained model downloads are automatic or require separate setup steps
- Memory and disk requirements for downloading and caching pretrained models
- Performance characteristics on single-machine vs. true distributed clusters
- Exact Python version support (classifiers list 3.6–3.9 but requires_python is unspecified)
Package facts
| License | permissive license permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 51 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,180,010 / month, #4,260 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: OS IndependentOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: Build ToolsTopic :: Software Development :: InternationalizationTopic :: Software Development :: Libraries :: Python ModulesTopic :: Software Development :: LocalizationTopic :: Text Processing :: LinguisticTyping :: Typed |
Evidence: spark_nlp-6.4.2-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “distributed NLP on Spark”
- spark-nlpSpark NLP provides distributed natural language processing on Apache…
- koalasKoalas implements the pandas DataFrame API on top of Apache Spark,…
- pyspark-pandasProvides tools for distributing Pandas DataFrames and Series across…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also flair · torchtext · spacy · textblob · keras-nlp · polyglot · google-cloud-language · urduhack · mleap · synapseml