minisbd
Free and open source library for fast sentence boundary detection
Decision gist · record as of 2026-08-14
Yes, if you need fast multilingual sentence segmentation and can accept AGPLv3 licensing. The package is lightweight, actively maintained, has no known vulnerabilities, and supports modern Python versions. Install friction is low. Main caveat: AGPLv3 requires source-code sharing for network deployments; review your use case before adopting in proprietary services.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- onnxruntime must be installed; optionally onnxruntime-gpu for GPU acceleration.
- Models are downloaded on first use to ~/.cache/minisbd.
- Low install friction: pure Python wheel with only three runtime dependencies (filelock, numpy, onnxruntime).
License · maintenance · safety
(agpl) — AGPLv3 copyleft: you must share source code of any modifications and provide access to the modified version if you operate it as a network service. Suitable for internal use and open-source projects; review required for proprietary deployments.
last release 2026-03-03 (164 days) · last repo commit 2026-03-23 · 8 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 245,515 downloads/mo, #8,730 on PyPI
Alternatives
Verify before relying
pip install minisbd
from minisbd import SBDetect
detector = SBDetect("en")
for sent in detector.sentences("Hello world. How are you?"):
print(sent)- Accuracy and performance benchmarks compared to other SBD libraries on multilingual corpora.
- Whether GPU acceleration via onnxruntime-gpu is automatically detected or requires explicit configuration.
- Behavior on edge cases: abbreviations, URLs, mixed-script text, and non-standard punctuation.
What it is and what it does
MiniSBD is a Python library for sentence boundary detection that splits raw text into individual sentences. It uses quantized ONNX models—the same models as Stanza but converted to ONNX format and compressed—to achieve fast, lightweight inference without heavy dependencies. The library supports a large set of languages and can run on CPU or GPU via onnxruntime.
You instantiate a detector with a language code (e.g., "en", "fr", "ja"), then call its sentences() method on text to get an iterable of detected sentences. Models are cached locally after first download. You can also supply a custom ONNX model path or change the cache directory at runtime. The package is free and open source under AGPLv3.
Use it for
- Preprocessing multilingual text corpora for NLP pipelines that require sentence-level input.
- Splitting user-submitted text into sentences for downstream tasks like translation, sentiment analysis, or named-entity recognition.
- Building text processing workflows that need language-aware sentence segmentation without heavy ML frameworks.
- Tokenizing documents in low-resource or embedded environments where model size and speed matter.
- Handling historical or non-standard text (Ancient Greek, Old French, Sanskrit) where rule-based splitters fail.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need fast multilingual sentence segmentation and can accept AGPLv3 licensing.
The package is lightweight, actively maintained, has no known vulnerabilities, and supports modern Python versions. Install friction is low. Main caveat: AGPLv3 requires source-code sharing for network deployments; review your use case before adopting in proprietary services.
Install
minisbd on PyPI
Before you install
Low install friction: pure Python wheel with only three runtime dependencies (filelock, numpy, onnxruntime). Active maintenance; last commit 2026-03-23.
onnxruntime must be installed; optionally onnxruntime-gpu for GPU acceleration. Models are downloaded on first use to ~/.cache/minisbd.
License in practice
AGPLv3 copyleft: you must share source code of any modifications and provide access to the modified version if you operate it as a network service. Suitable for internal use and open-source projects; review required for proprietary deployments.
Quickstart
pip install minisbd
from minisbd import SBDetect
detector = SBDetect("en")
for sent in detector.sentences("Hello world. How are you?"):
print(sent)
Verify before relying
- Accuracy and performance benchmarks compared to other SBD libraries on multilingual corpora.
- Whether GPU acceleration via onnxruntime-gpu is automatically detected or requires explicit configuration.
- Behavior on edge cases: abbreviations, URLs, mixed-script text, and non-standard punctuation.
Package facts
| License | Not declared agpl |
| Python support | Supports the current Python release >=3.6 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesfilelocknumpyonnxruntime |
| Maintenance | Actively maintained 164 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 245,515 / month, #8,730 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: GNU Affero General Public License v3Operating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.6Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: minisbd-0.9.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “multilingual sentence splitting”
- minisbdDetects sentence boundaries in text across many languages using…
- simplemmaSimplemma converts inflected word forms to their dictionary base…
- segtokSplits Indo-European text into sentences and words using rule-based…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also segtok · pysbd · sentence-stream · jieba3k · semantic-text-splitter · fasttext-langdetect · jieba · razdel · cnstd · simplemma