Sastrawi
Library for stemming Indonesian (Bahasa) text
Decision gist · record as of 2026-08-14
Yes, if you need Indonesian stemming. The package is stable, has no dependencies, and carries a permissive license. The 2016 release date is a concern for Python version compatibility—verify it works with your target Python version before relying on it in production.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction install with no runtime dependencies.
- Last release was in 2016, but the repository remains active with recent commits and is marked production-stable.
License · maintenance · safety
MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
last release 2016-01-18 (3861 days) · last repo commit 2026-05-21 · 361 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 79,895 downloads/mo, #14,330 on PyPI
Alternatives
Verify before relying
pip install Sastrawi
from Sastrawi.Stemmer.StemmerFactory import StemmerFactory
factory = StemmerFactory()
stemmer = factory.create_stemmer()
output = stemmer.stem('Perekonomian Indonesia sedang dalam pertumbuhan')
print(output) # ekonomi indonesia sedang dalam tumbuh- Whether the package supports modern Python versions beyond what was current in 2016
- Performance characteristics on large Indonesian text corpora
- Accuracy of stemming against standard Indonesian linguistic benchmarks
What it is and what it does
Sastrawi is a Python port of the original PHP Sastrawi project for Indonesian language stemming. It takes inflected Indonesian words and reduces them to their root form by removing affixes and morphological markers, which is a common preprocessing step in Indonesian text analysis and natural language processing tasks.
The package provides a factory-based API: you instantiate a StemmerFactory, create a stemmer instance, and call its stem() method on text. It handles both single words and full sentences, stripping prefixes, suffixes, and other morphological variations specific to Indonesian grammar. With no external runtime dependencies and a permissive MIT license, it integrates easily into Indonesian NLP pipelines.
Use it for
- Preprocess Indonesian text for search engines or information retrieval systems to normalize word variations
- Reduce Indonesian documents to root forms before feeding them into machine learning classifiers or topic models
- Build Indonesian language chatbots or question-answering systems that need to match user queries to base word forms
- Analyze Indonesian social media or news text by normalizing inflected words for frequency analysis or sentiment detection
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need Indonesian stemming.
The package is stable, has no dependencies, and carries a permissive license. The 2016 release date is a concern for Python version compatibility—verify it works with your target Python version before relying on it in production.
Install
sastrawi on PyPI
Before you install
Low friction install with no runtime dependencies. Last release was in 2016, but the repository remains active with recent commits and is marked production-stable.
License in practice
MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
Quickstart
pip install Sastrawi
from Sastrawi.Stemmer.StemmerFactory import StemmerFactory
factory = StemmerFactory()
stemmer = factory.create_stemmer()
output = stemmer.stem('Perekonomian Indonesia sedang dalam pertumbuhan')
print(output) # ekonomi indonesia sedang dalam tumbuh
Verify before relying
- Whether the package supports modern Python versions beyond what was current in 2016
- Performance characteristics on large Indonesian text corpora
- Accuracy of stemming against standard Indonesian linguistic benchmarks
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 3,861 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 79,895 / month, #14,330 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/Stable |
Evidence: Sastrawi-1.0.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “indonesian text stemming”
- SastrawiSastrawi reduces inflected Indonesian words to their base form (stem)…
- PyStemmerPyStemmer provides word stemming algorithms for multiple languages,…
- snowballstemmerProvides stemming algorithms for 34 languages, reducing word variants…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also PyStemmer · snowballstemmer · sea-g2p · kiwipiepy-model · jieba · py-rust-stemmers · lemminflect · kiwipiepy · stem · wordninja