$npx skillfedfor your agent

semantic-text-splitter

Split text into semantic chunks, up to a desired chunk size. Supports calculating length by characters and tokens, and is callable from Rust and Python.

With conditionsPyPI LinguisticReleased Jun 2026307.4K downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — semantic_text_splitter-0.32.0-cp310-abi3-macosx_10_12_x86_64.whl · semantic_text_splitter-0.32.0-cp310-abi3-macosx_11_0_arm64.whl · semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_aarch64.whl
v0.32.0 · released 2026-06-16 · Python >=3.10

Yes, if you need semantic-aware text chunking for LLM workflows. The library is actively maintained, has no known vulnerabilities, and offers a cleaner API than character-only splitting. The MIT license poses no restrictions. Medium install friction (Rust compilation) is a minor trade-off for the performance and semantic quality it provides. Install if you're building RAG systems, prompt pipelines, or document processing for language models.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; compiled Rust wheels may not be available for all architectures.
  • Medium install friction due to compiled Rust bindings, but pre-built wheels cover common platforms (x86_64, ARM, Windows).
  • Requires Python 3.10+.

License · maintenance · safety

MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with only attribution required—no restrictions on commercial or proprietary use.

last release 2026-06-16 (59 days) · last repo commit 2026-08-14 · 625 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 307,442 downloads/mo, #7,775 on PyPI

Verify before relying

from semantic_text_splitter import TextSplitter

splitter = TextSplitter(max_characters=1000)
chunks = splitter.chunks("your document text")
  • Performance characteristics (speed, memory overhead) compared to alternatives like LangChain's TextSplitter.
  • Whether custom tokenizers beyond Hugging Face and Tiktoken are supported.
  • Behavior when a single semantic unit exceeds the specified chunk size limit.
Same gist for agents: .md · .json

What it is and what it does

semantic-text-splitter is a Python library that breaks long documents into smaller chunks optimized for LLM processing. Rather than splitting at fixed character boundaries, it respects semantic structure—sentences, paragraphs, markdown blocks, and newline sequences—to keep related content together. You can specify chunk size by character count, token range, or custom callback, and it supports multiple tokenizer backends (Hugging Face, Tiktoken) or plain character counting.

The library provides two main splitters: TextSplitter for plain text and MarkdownSplitter for markdown documents. It uses a hierarchical approach that tries to fill chunks to your target size while never breaking at lower semantic levels if a higher-level boundary is available. This is useful when preparing documents for RAG pipelines, prompt engineering, or any workflow where you need to feed text to models with fixed context limits while preserving meaning.

Use it for

  • Prepare long documents for retrieval-augmented generation (RAG) by splitting into token-bounded chunks that respect paragraph structure.
  • Split markdown documentation into semantic sections for indexing and search without breaking code blocks or inline formatting.
  • Chunk text for fine-tuning datasets where you need to respect sentence and paragraph boundaries to preserve training signal.
  • Prepare long articles or books for LLM summarization by splitting into context-window-sized pieces that maintain narrative coherence.
  • Build a document ingestion pipeline that respects both character/token limits and semantic structure for downstream NLP tasks.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need semantic-aware text chunking for LLM workflows.

The library is actively maintained, has no known vulnerabilities, and offers a cleaner API than character-only splitting. The MIT license poses no restrictions. Medium install friction (Rust compilation) is a minor trade-off for the performance and semantic quality it provides. Install if you're building RAG systems, prompt pipelines, or document processing for language models.

Install

semantic-text-splitter on PyPI

Before you install

Medium install friction due to compiled Rust bindings, but pre-built wheels cover common platforms (x86_64, ARM, Windows). Requires Python 3.10+. Active maintenance with recent releases.

Requires Python 3.10 or later; compiled Rust wheels may not be available for all architectures.

License in practice

MIT license permits unrestricted use, modification, and distribution with only attribution required—no restrictions on commercial or proprietary use.

Quickstart

from semantic_text_splitter import TextSplitter

splitter = TextSplitter(max_characters=1000)
chunks = splitter.chunks("your document text")

Verify before relying

  • Performance characteristics (speed, memory overhead) compared to alternatives like LangChain's TextSplitter.
  • Whether custom tokenizers beyond Hugging Face and Tiktoken are supported.
  • Behavior when a single semantic unit exceeds the specified chunk size limit.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 59 days since the last release
Last repo commit
First released
Downloads307,442 / month, #7,775 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyProgramming Language :: Rust

Evidence: semantic_text_splitter-0.32.0-cp310-abi3-macosx_10_12_x86_64.whl; semantic_text_splitter-0.32.0-cp310-abi3-macosx_11_0_arm64.whl; semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_aarch64.whl; semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_armv7l.whl; semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_ppc64le.whl; semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_s390x.whl; semantic_text_splitter-0.32.0-cp310-abi3-manylinux_2_28_x86_64.whl; semantic_text_splitter-0.32.0-cp310-abi3-win32.whl; semantic_text_splitter-0.32.0-cp310-abi3-win_amd64.whl; semantic_text_splitter-0.32.0-cp314-cp314t-macosx_10_12_x86_64.whl; semantic_text_splitter-0.32.0-cp314-cp314t-macosx_11_0_arm64.whl; semantic_text_splitter-0.32.0-cp314-cp314t-manylinux_2_28_aarch64.whl; semantic_text_splitter-0.32.0-cp314-cp314t-manylinux_2_28_armv7l.whl; semantic_text_splitter-0.32.0-cp314-cp314t-manylinux_2_28_ppc64le.whl; semantic_text_splitter-0.32.0-cp314-cp314t-manylinux_2_28_s390x.whl; semantic_text_splitter-0.32.0-cp314-cp314t-manylinux_2_28_x86_64.whl; semantic_text_splitter-0.32.0-cp314-cp314t-win32.whl; semantic_text_splitter-0.32.0-cp314-cp314t-win_amd64.whl; semantic_text_splitter-0.32.0-cp314-cp314-win32.whl; semantic_text_splitter-0.32.0-cp314-cp314-win_amd64.whl

Tags

Capabilities
text chunking for llm contextsemantic text splittingdocument chunking by tokenssplit text into chunksnlp text segmentationmarkdown document splittingtokenizer-aware text splitting
Topics
llm-toolingdocument-processingtokenization
PyPI keywords
textsplittokenizernlpai

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “text chunking for llm context”

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also chonkie-core · sentence-stream · semchunk · memchunk · chonkie · uniseg · langchain-text-splitters · minisbd · sumy · paragraphs

Further reading