$npx skillfedfor your agent

simplemma

Fast and zero-dependency lemmatization, tokenization and sentence splitting for 54 languages.

Worth itPyPI Information AnalysisReleased Aug 2026220.5K downloads / moMIT LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — simplemma-2.0.0-py3-none-any.whl
v2.0.0 · released 2026-08-12 · Python >=3.10

Yes. Simplemma is worth installing if you need fast, offline lemmatization across many languages without external dependencies or model downloads. Its zero-dependency design, active maintenance, MIT license, and absence of known vulnerabilities make it a low-risk choice. Install it for baseline NLP work, teaching, or low-resource settings; do not install it if you need the highest accuracy and can afford the overhead of neural pipelines.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; simplemma==1.1.2 is the last version supporting 3.8 and 3.9.
  • Installation is straightforward with zero runtime dependencies and a 19 MB footprint.
  • The package is actively maintained with a recent release (2 days old) and no known vulnerabilities.

License · maintenance · safety

MIT License (permissive) — MIT License permits unrestricted use, modification, and distribution in both open-source and commercial contexts with minimal obligations.

last release 2026-08-12 (2 days) · last repo commit 2026-08-12 · 215 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 220,470 downloads/mo, #9,299 on PyPI

Verify before relying

pip install simplemma

import simplemma

simplemma.lemmatize('masks', lang='en')
# 'mask'

simplemma.text_lemmatizer('Hier sind Vaccines.', lang=('de', 'en'))
# ['hier', 'sein', 'vaccine', '.']
  • Whether accuracy figures (0.91–0.97 for 34 languages, 0.85–0.90 for morphologically rich ones) remain current across all 54 languages in version 2.0.0.
  • Whether the ~1.9M tokens/s (German) and ~3.4M (English) throughput claims hold in real-world workloads outside the benchmarks cited.
  • Whether optional marisa-trie dependency (for lowest memory usage) is compatible with all 54 languages or has language-specific limitations.
Same gist for agents: .md · .json

What it is and what it does

Simplemma is a pure-Python lemmatizer that reduces inflected word forms to their base dictionary forms across 54 languages. It works offline with no model downloads, ships in a 19 MB package, and includes built-in tokenization, sentence splitting, and language detection utilities. The core trade-off is deliberate: it sacrifices the accuracy of neural pipelines (typically a few percentage points behind trained models) in exchange for speed (millions of tokens per second), simplicity, and a small footprint that suits low-resource settings, teaching, and baseline NLP work.

Unlike stemming, lemmatization always returns valid linguistic forms. Simplemma handles this without morphosyntactic information by performing dictionary lookups on raw token sequences. It supports language chaining to improve coverage when a word is unknown in one language but recognized in another, and offers a tunable RAM footprint via a `low_memory` flag or optional marisa-trie backend. The package is actively maintained, has no runtime dependencies, and works on current Python versions (3.10+).

Use it for

  • Build a search engine or text indexing system where lemmatization reduces vocabulary size without downloading language models.
  • Preprocess multilingual documents for topic modeling or information retrieval in resource-constrained environments.
  • Teach NLP fundamentals in a classroom setting where a simple, dependency-free tool is easier to install and explain than neural pipelines.
  • Establish a baseline lemmatization result to compare against more complex models in morphological analysis research.
  • Detect the language of short text snippets by scoring them against a set of candidate languages.
  • Tokenize and split sentences in a rule-based pipeline where speed and determinism matter more than state-of-the-art accuracy.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Simplemma is worth installing if you need fast, offline lemmatization across many languages without external dependencies or model downloads. Its zero-dependency design, active maintenance, MIT license, and absence of known vulnerabilities make it a low-risk choice. Install it for baseline NLP work, teaching, or low-resource settings; do not install it if you need the highest accuracy and can afford the overhead of neural pipelines.

Install

simplemma on PyPI

Before you install

Installation is straightforward with zero runtime dependencies and a 19 MB footprint. The package is actively maintained with a recent release (2 days old) and no known vulnerabilities.

Requires Python 3.10 or later; simplemma==1.1.2 is the last version supporting 3.8 and 3.9.

License in practice

MIT License permits unrestricted use, modification, and distribution in both open-source and commercial contexts with minimal obligations.

Quickstart

pip install simplemma

import simplemma

simplemma.lemmatize('masks', lang='en')
# 'mask'

simplemma.text_lemmatizer('Hier sind Vaccines.', lang=('de', 'en'))
# ['hier', 'sein', 'vaccine', '.']

Verify before relying

  • Whether accuracy figures (0.91–0.97 for 34 languages, 0.85–0.90 for morphologically rich ones) remain current across all 54 languages in version 2.0.0.
  • Whether the ~1.9M tokens/s (German) and ~3.4M (English) throughput claims hold in real-world workloads outside the benchmarks cited.
  • Whether optional marisa-trie dependency (for lowest memory usage) is compatible with all 54 languages or has language-specific limitations.

Package facts

LicenseMIT License permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads220,470 / month, #9,299 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseNatural Language :: ArabicNatural Language :: ArmenianNatural Language :: BosnianNatural Language :: BulgarianNatural Language :: CatalanNatural Language :: CroatianNatural Language :: CzechNatural Language :: DanishNatural Language :: DutchNatural Language :: EnglishNatural Language :: EsperantoNatural Language :: EstonianNatural Language :: FinnishNatural Language :: FrenchNatural Language :: GalicianNatural Language :: GeorgianNatural Language :: GermanNatural Language :: GreekNatural Language :: HebrewNatural Language :: HindiNatural Language :: HungarianNatural Language :: IcelandicNatural Language :: IndonesianNatural Language :: IrishNatural Language :: ItalianNatural Language :: LatinNatural Language :: LatvianNatural Language :: LithuanianNatural Language :: MacedonianNatural Language :: MalayNatural Language :: NorwegianNatural Language :: PersianNatural Language :: PolishNatural Language :: PortugueseNatural Language :: RomanianNatural Language :: RussianNatural Language :: SerbianNatural Language :: SlovakNatural Language :: SlovenianNatural Language :: SpanishNatural Language :: SwedishNatural Language :: TurkishNatural Language :: UkrainianOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.15Topic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: InternationalizationTopic :: Software Development :: LocalizationTopic :: Text Processing :: LinguisticTyping :: Typed

Evidence: simplemma-2.0.0-py3-none-any.whl

Tags

Capabilities
multilingual lemmatizationlemmatizer no dependenciesfast lemmatization pythonoffline lemmatizationtokenization sentence splittinglanguage detectionmorphological analysis
Topics
multilingualoffline-firstzero-dependencies
PyPI keywords
language detectionlanguage identificationlangidlemmatizationlemmatizerlemmatisernlpsentence segmentationsentence splittingtokenizationtokenizer

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “multilingual lemmatization”

  • simplemmaSimplemma converts inflected word forms to their dictionary base…
  • wnWn is a Python library for querying and exploring wordnet…
  • yakeYAKE extracts keywords from text documents using unsupervised…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also lemminflect · udapi · pymorphy3-dicts-uk · pymorphy3 · pymorphy3-dicts-ru · minisbd · polyglot · jieba · num2words · wn