$npx skillfedfor your agent

Sastrawi

Library for stemming Indonesian (Bahasa) text

With conditionsPyPI LinguisticReleased Jan 201679.9K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — Sastrawi-1.0.1-py2.py3-none-any.whl
v1.0.1 · released 2016-01-18

Yes, if you need Indonesian stemming. The package is stable, has no dependencies, and carries a permissive license. The 2016 release date is a concern for Python version compatibility—verify it works with your target Python version before relying on it in production.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low friction install with no runtime dependencies.
  • Last release was in 2016, but the repository remains active with recent commits and is marked production-stable.

License · maintenance · safety

MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

last release 2016-01-18 (3861 days) · last repo commit 2026-05-21 · 361 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 79,895 downloads/mo, #14,330 on PyPI

Verify before relying

pip install Sastrawi

from Sastrawi.Stemmer.StemmerFactory import StemmerFactory
factory = StemmerFactory()
stemmer = factory.create_stemmer()
output = stemmer.stem('Perekonomian Indonesia sedang dalam pertumbuhan')
print(output)  # ekonomi indonesia sedang dalam tumbuh
  • Whether the package supports modern Python versions beyond what was current in 2016
  • Performance characteristics on large Indonesian text corpora
  • Accuracy of stemming against standard Indonesian linguistic benchmarks
Same gist for agents: .md · .json

What it is and what it does

Sastrawi is a Python port of the original PHP Sastrawi project for Indonesian language stemming. It takes inflected Indonesian words and reduces them to their root form by removing affixes and morphological markers, which is a common preprocessing step in Indonesian text analysis and natural language processing tasks.

The package provides a factory-based API: you instantiate a StemmerFactory, create a stemmer instance, and call its stem() method on text. It handles both single words and full sentences, stripping prefixes, suffixes, and other morphological variations specific to Indonesian grammar. With no external runtime dependencies and a permissive MIT license, it integrates easily into Indonesian NLP pipelines.

Use it for

  • Preprocess Indonesian text for search engines or information retrieval systems to normalize word variations
  • Reduce Indonesian documents to root forms before feeding them into machine learning classifiers or topic models
  • Build Indonesian language chatbots or question-answering systems that need to match user queries to base word forms
  • Analyze Indonesian social media or news text by normalizing inflected words for frequency analysis or sentiment detection

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need Indonesian stemming.

The package is stable, has no dependencies, and carries a permissive license. The 2016 release date is a concern for Python version compatibility—verify it works with your target Python version before relying on it in production.

Install

sastrawi on PyPI

Before you install

Low friction install with no runtime dependencies. Last release was in 2016, but the repository remains active with recent commits and is marked production-stable.

License in practice

MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.

Quickstart

pip install Sastrawi

from Sastrawi.Stemmer.StemmerFactory import StemmerFactory
factory = StemmerFactory()
stemmer = factory.create_stemmer()
output = stemmer.stem('Perekonomian Indonesia sedang dalam pertumbuhan')
print(output)  # ekonomi indonesia sedang dalam tumbuh

Verify before relying

  • Whether the package supports modern Python versions beyond what was current in 2016
  • Performance characteristics on large Indonesian text corpora
  • Accuracy of stemming against standard Indonesian linguistic benchmarks

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependenciesNone
MaintenanceActively maintained 3,861 days since the last release
Last repo commit
First released
Downloads79,895 / month, #14,330 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/Stable

Evidence: Sastrawi-1.0.1-py2.py3-none-any.whl

Tags

Capabilities
indonesian text stemmingbahasa indonesia stemmerindonesian language processingmorphological analysis indonesianindonesian nlp stemmingreduce indonesian word formsindonesian linguistic preprocessing
Topics
indonesian-nlpmorphological-analysis
PyPI keywords
linguisticstemmingindonesianbahasa

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “indonesian text stemming”

  • SastrawiSastrawi reduces inflected Indonesian words to their base form (stem)…
  • PyStemmerPyStemmer provides word stemming algorithms for multiple languages,…
  • snowballstemmerProvides stemming algorithms for 34 languages, reducing word variants…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also PyStemmer · snowballstemmer · sea-g2p · kiwipiepy-model · jieba · py-rust-stemmers · lemminflect · kiwipiepy · stem · wordninja