$npx skillfedfor your agent

sentencepiece

Unsupervised text tokenizer and detokenizer.

Worth itPyPI Python ModulesReleased Jul 202635.6M downloads / moApache-2.0Platform wheel

Decision gist · record as of 2026-08-14

platform wheels — sentencepiece-0.2.2-cp310-cp310-macosx_10_9_universal2.whl · sentencepiece-0.2.2-cp310-cp310-macosx_10_9_x86_64.whl · sentencepiece-0.2.2-cp310-cp310-macosx_11_0_arm64.whl
v0.2.2 · released 2026-07-12 · Python >=3.9

Yes. SentencePiece is a stable, widely-used tokenizer with no known vulnerabilities, permissive licensing, and active maintenance. Install friction is moderate but manageable due to prebuilt wheels. It is essential for projects requiring standardized multilingual tokenization or integration with models trained on SentencePiece vocabularies.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a pre-trained SentencePiece model file (.model) to load before encoding or decoding.
  • Medium install friction due to compiled wheels for multiple platforms and Python versions (3.9–3.14).
  • Prebuilt wheels are available for Linux (x86_64/aarch64), macOS, and Windows (x64/arm64), reducing build complexity.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing use in commercial and proprietary projects with minimal restrictions beyond attribution and liability disclaimers.

last release 2026-07-12 (33 days) · last repo commit 2026-08-14 · 12,022 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 35,582,256 downloads/mo, #744 on PyPI

Verify before relying

import sentencepiece as spm

sp = spm.SentencePieceProcessor(model_file='model.model')
ids = sp.encode('This is a test')
text = sp.decode([284, 47, 11, 4, 15, 400])
  • Whether training new SentencePiece models is supported through this Python wrapper or only inference.
  • Performance characteristics for very large batch sizes or high-throughput scenarios.
  • Memory overhead when decoding large token sequences or handling very long texts.
Same gist for agents: .md · .json

What it is and what it does

SentencePiece is a Python wrapper around Google's SentencePiece C++ library for unsupervised text tokenization. It breaks text into subword units (tokens) and converts them to integer IDs, or reverses the process to reconstruct text from token sequences. The library supports polymorphic inputs—single strings or batches—and offers multiple output formats: integer IDs, string pieces, NumPy arrays, byte-level representations, offset mappings that link tokens back to character positions in the original text, and Protobuf messages.

The package is designed for NLP pipelines where consistent, language-agnostic tokenization is needed. It handles Unicode text and raw bytes, supports sampling-based encoding for subword regularization, and can generate n-best alternative tokenizations. Batch operations automatically release Python's GIL and run in parallel in C++, making it suitable for processing large document collections. No runtime Python dependencies are required—the wrapper is a thin layer over precompiled C++ binaries.

Use it for

  • Tokenizing multilingual text for transformer models or other NLP systems that require consistent subword segmentation.
  • Converting token IDs back to readable text in inference pipelines or when debugging model outputs.
  • Mapping tokens to character offsets in the original text for tasks like named-entity recognition or text highlighting.
  • Batch processing large document collections with parallel C++ execution to reduce tokenization latency.
  • Applying subword regularization during training by sampling alternative tokenizations of the same text.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

SentencePiece is a stable, widely-used tokenizer with no known vulnerabilities, permissive licensing, and active maintenance. Install friction is moderate but manageable due to prebuilt wheels. It is essential for projects requiring standardized multilingual tokenization or integration with models trained on SentencePiece vocabularies.

Install

sentencepiece on PyPI

Before you install

Medium install friction due to compiled wheels for multiple platforms and Python versions (3.9–3.14). Prebuilt wheels are available for Linux (x86_64/aarch64), macOS, and Windows (x64/arm64), reducing build complexity. Active maintenance with recent releases.

Requires a pre-trained SentencePiece model file (.model) to load before encoding or decoding.

License in practice

Licensed under Apache-2.0 (permissive), allowing use in commercial and proprietary projects with minimal restrictions beyond attribution and liability disclaimers.

Quickstart

import sentencepiece as spm

sp = spm.SentencePieceProcessor(model_file='model.model')
ids = sp.encode('This is a test')
text = sp.decode([284, 47, 11, 4, 15, 400])

Verify before relying

  • Whether training new SentencePiece models is supported through this Python wrapper or only inference.
  • Performance characteristics for very large batch sizes or high-throughput scenarios.
  • Memory overhead when decoding large token sequences or handling very long texts.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 33 days since the last release
Last repo commit
First released
Downloads35,582,256 / month, #744 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9Programming Language :: Python :: Free Threading :: 2 - BetaTopic :: Software Development :: Libraries :: Python ModulesTopic :: Text Processing :: Linguistic

Evidence: sentencepiece-0.2.2-cp310-cp310-macosx_10_9_universal2.whl; sentencepiece-0.2.2-cp310-cp310-macosx_10_9_x86_64.whl; sentencepiece-0.2.2-cp310-cp310-macosx_11_0_arm64.whl; sentencepiece-0.2.2-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; sentencepiece-0.2.2-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; sentencepiece-0.2.2-cp310-cp310-win_amd64.whl; sentencepiece-0.2.2-cp310-cp310-win_arm64.whl; sentencepiece-0.2.2-cp311-cp311-macosx_10_9_universal2.whl; sentencepiece-0.2.2-cp311-cp311-macosx_10_9_x86_64.whl; sentencepiece-0.2.2-cp311-cp311-macosx_11_0_arm64.whl; sentencepiece-0.2.2-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; sentencepiece-0.2.2-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; sentencepiece-0.2.2-cp311-cp311-win_amd64.whl; sentencepiece-0.2.2-cp311-cp311-win_arm64.whl; sentencepiece-0.2.2-cp312-cp312-macosx_10_13_universal2.whl; sentencepiece-0.2.2-cp312-cp312-macosx_10_13_x86_64.whl; sentencepiece-0.2.2-cp312-cp312-macosx_11_0_arm64.whl; sentencepiece-0.2.2-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; sentencepiece-0.2.2-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; sentencepiece-0.2.2-cp312-cp312-win_amd64.whl

Tags

Capabilities
text tokenization encoding decodingsubword tokenizersentencepiece tokenizernlp text segmentationtoken id conversionbatch text processingunicode character offsets
Topics
tokenizationnlpmultilingual

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “text tokenization encoding decoding”

  • sentencepieceSentencePiece is an unsupervised text tokenizer and detokenizer that…
  • tokieA fast, Rust-backed tokenizer library that encodes and decodes text…
  • tiktokentiktoken is a fast BPE tokenizer that converts text into token…

Give your agent the search over MCP, or paste the wish link into any chat.

More Python Modules packages

idna Worth it
PyPI · Python Modules · released Jun 2026

Converts domain names between Unicode and ASCII-compatible encoding (Punycode) according to IDNA 2008 and Unicode Technical Standard 46, with security validation and broader script coverage than the standard library.

Install it if you work with internationalized domain names, need to validate domains, or use HTTP clients that depend on it transitively.

BSD-3-Clausepure Python · 3.9+
1.8Bdownloads / mo
setuptools Worth it
PyPI · Python Modules · released Aug 2026

Setuptools is a Python build backend and package management tool that handles building, distributing, and installing Python packages, including support for C/C++ extension modules.

MITpure Python · 3.10+
1.6Bdownloads / mo
PyYAML Worth it
PyPI · Python Modules · released Sep 2025

PyYAML parses and emits YAML 1.1 data format, enabling serialization and deserialization of configuration files and Python objects to and from human-readable YAML text.

MITcompiled wheel · 3.8+
1.2Bdownloads / mo
pydantic Worth it
PyPI · Python Modules · released May 2026

Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.

MITpure Python · 3.9+
1.1Bdownloads / mo
annotated-types Worth it
PyPI · Python Modules · released Jul 2026

Provides reusable metadata objects for use with PEP-593 `typing.Annotated` to express common constraints like bounds, collection sizes, and predicates on types.

Install it if you use or build libraries that need to express type constraints in a standardized, inspectable way—or if you want to annotate your own types with…

MITpure Python · 3.10+
871.3Mdownloads / mo
typing-inspection Worth it
PyPI · Python Modules · released Aug 2026

Provides runtime tools to inspect and introspect Python type annotations, enabling programmatic examination of type hints at execution time.

MITpure Python · 3.10+
783.0Mdownloads / mo

See also curated-tokenizers · segments · jieba3k · tensorflow-text · tokie · konoha · semchunk · blingfire · jieba · pytorch-tokenizers