$npx skillfedfor your agent

tinysegmenter

Very compact Japanese tokenizer

SkipPyPI Artificial IntelligenceReleased Sep 2018510.1K downloads / moNew BSDSource build

Decision gist · record as of 2026-08-14

sdist only — tinysegmenter-0.4.tar.gz · builds from source
v0.4 · released 2018-09-16

No, unless you have a specific, narrow need for a lightweight Japanese tokenizer and can tolerate an abandoned codebase. The package has not been maintained since 2018, and there is no guarantee it will work correctly with modern Python versions or character sets. For active projects, consider a maintained alternative like MeCab, Janome, or a modern NLP library.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 2.6 or above (including Python 3); no other external dependencies, but the package is unmaintained since 2018.
  • High install friction: the package is abandoned (last release 2018-09-16, 2889 days ago) with no active maintenance.
  • No runtime dependencies, but the age and lack of ongoing support mean you are on your own for any issues.

License · maintenance · safety

New BSD (permissive) — Distributed under New BSD License (permissive), so you can use, modify, and redistribute freely with attribution and no warranty.

last release 2018-09-16 (2889 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 510,066 downloads/mo, #6,269 on PyPI

Verify before relying

import tinysegmenter
segmenter = tinysegmenter.TinySegmenter()
tokens = segmenter.tokenize(u"私の名前は中野です")
print(' | '.join(tokens))
  • Whether the tokenizer's accuracy and behavior remain adequate for modern Japanese text and character sets.
  • Compatibility with Python versions beyond what was tested at the time of the last release.
  • Whether the package works correctly with contemporary dependency versions if you need to integrate it into a larger stack.
Same gist for agents: .md · .json

What it is and what it does

TinySegmenter is a Python port of a compact Japanese tokenizer originally written in JavaScript. It segments Japanese text into individual morphological units (words, particles, suffixes) without requiring a dictionary or machine learning model, making it extremely lightweight and fast. The library exposes a simple API: instantiate a TinySegmenter object and call its tokenize() method on a Unicode string to get a list of tokens.

The package is designed for straightforward Japanese text processing tasks where you need basic word segmentation without the overhead of larger NLP frameworks. It has no runtime dependencies and works as a standalone module. However, the project is abandoned—the last release was in 2018 and there is no active maintenance, so you should expect no bug fixes or updates.

Use it for

  • Segment Japanese text into tokens for search indexing or keyword extraction in a lightweight application.
  • Tokenize Japanese user input for simple text analysis or filtering without pulling in heavy NLP dependencies.
  • Use as a baseline tokenizer in NLTK pipelines by subclassing both TinySegmenter and NLTK's TokenizerI interface.
  • Process Japanese text in resource-constrained environments where dictionary-based or neural tokenizers are too heavy.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Skip

No, unless you have a specific, narrow need for a lightweight Japanese tokenizer and can tolerate an abandoned codebase.

The package has not been maintained since 2018, and there is no guarantee it will work correctly with modern Python versions or character sets. For active projects, consider a maintained alternative like MeCab, Janome, or a modern NLP library.

Install

tinysegmenter on PyPI

Before you install

High install friction: the package is abandoned (last release 2018-09-16, 2889 days ago) with no active maintenance. No runtime dependencies, but the age and lack of ongoing support mean you are on your own for any issues.

Requires Python 2.6 or above (including Python 3); no other external dependencies, but the package is unmaintained since 2018.

License in practice

Distributed under New BSD License (permissive), so you can use, modify, and redistribute freely with attribution and no warranty.

Quickstart

import tinysegmenter
segmenter = tinysegmenter.TinySegmenter()
tokens = segmenter.tokenize(u"私の名前は中野です")
print(' | '.join(tokens))

Verify before relying

  • Whether the tokenizer's accuracy and behavior remain adequate for modern Japanese text and character sets.
  • Compatibility with Python versions beyond what was tested at the time of the last release.
  • Whether the package works correctly with contemporary dependency versions if you need to integrate it into a larger stack.

Package facts

LicenseNew BSD permissive
Python supportNot specified
Install frictionHigh. Source build required
Runtime dependenciesNone
MaintenanceAbandoned 2,889 days since the last release
First released
Downloads510,066 / month, #6,269 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaLicense :: OSI Approved :: BSD LicenseOperating System :: POSIX :: LinuxProgramming Language :: PythonTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text Processing :: Linguistic

Evidence: tinysegmenter-0.4.tar.gz

Tags

Capabilities
japanese text tokenizationjapanese morphological segmentationcompact japanese tokenizerjapanese word segmentationlightweight nlp japanesejapanese language processing
Topics
japanese-nlptokenizationabandoned

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “japanese morphological segmentation”

  • tinysegmenterTinySegmenter is a compact Japanese tokenizer that breaks Japanese…
  • SudachiPySudachiPy is a Python binding for Sudachi.rs, a Japanese…
  • JanomeJanome is a Japanese morphological analyzer written in pure Python…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also Janome · nagisa · segtok · mecab-python3 · jieba · jieba3k · mecab · SudachiDict-small · ja-ginza · SudachiPy

Further reading