memchunk
The fastest semantic text chunking library
Decision gist · record as of 2026-08-14
Yes, if you need fast semantic text chunking for RAG, NLP pipelines, or high-volume document processing. The library is actively maintained, has no known vulnerabilities, supports modern Python versions, and offers a permissive dual license. Install friction is moderate due to compiled wheels, but pre-built binaries are available for all major platforms. Not necessary if chunking is not a performance bottleneck in your workflow.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8 or later.
- Pre-compiled wheels are provided; source builds require a Rust toolchain.
- Medium install friction due to compiled wheels, but pre-built binaries are available for Python 3.8–3.14 across Linux, macOS (x86_64 and ARM), and Windows.
License · maintenance · safety
MIT OR Apache-2.0 (permissive) — Dual-licensed under MIT or Apache-2.0 (permissive); you may choose either license. No restrictions on commercial or proprietary use.
last release 2026-01-05 (221 days) · last repo commit 2026-05-28 · 358 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 96,649 downloads/mo, #13,201 on PyPI
Alternatives
Verify before relying
from memchunk import Chunker
text = "Hello world. How are you?"
for chunk in Chunker(text):
print(bytes(chunk))
# Custom size and delimiters
for chunk in Chunker(text, size=1024, delimiters=".?!\n"):
print(bytes(chunk))- Actual throughput benchmarks and whether '1 TB/s' claim is measured on specific hardware or representative of typical use.
- Memory overhead of the Chunker object and whether it scales linearly with text size.
- Behavior when text contains multi-byte UTF-8 sequences at chunk boundaries.
What it is and what it does
memchunk is a text chunking library optimized for speed using SIMD instructions and lookup tables. It splits text at semantic boundaries (periods, newlines, or custom delimiters) and returns chunks as zero-copy memoryview objects. The library is designed for high-throughput scenarios such as preparing documents for retrieval-augmented generation (RAG) pipelines or tokenization workflows.
The package provides a simple Chunker class that accepts text and configuration options: chunk size (default 4KB), delimiter characters, multi-byte patterns (useful for tokenizer-specific markers), and fallback strategies for finding split points. It supports Python 3.8 through 3.14 and is implemented in Rust with Python bindings, providing pre-compiled wheels for common platforms.
Use it for
- Prepare large document collections for RAG systems by splitting text into semantic chunks before embedding.
- Tokenize and chunk text for language model input pipelines where consistent chunk boundaries matter.
- Process streaming or batch text data at high throughput when chunking is a bottleneck.
- Split documents at custom delimiters (e.g., SentencePiece metaspace markers) for specialized NLP workflows.
- Reduce memory overhead in text processing by using zero-copy memoryview chunks instead of string copies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need fast semantic text chunking for RAG, NLP pipelines, or high-volume document processing.
The library is actively maintained, has no known vulnerabilities, supports modern Python versions, and offers a permissive dual license. Install friction is moderate due to compiled wheels, but pre-built binaries are available for all major platforms. Not necessary if chunking is not a performance bottleneck in your workflow.
Install
memchunk on PyPI
Before you install
Medium install friction due to compiled wheels, but pre-built binaries are available for Python 3.8–3.14 across Linux, macOS (x86_64 and ARM), and Windows. Active maintenance with recent releases.
Requires Python 3.8 or later. Pre-compiled wheels are provided; source builds require a Rust toolchain.
License in practice
Dual-licensed under MIT or Apache-2.0 (permissive); you may choose either license. No restrictions on commercial or proprietary use.
Quickstart
from memchunk import Chunker
text = "Hello world. How are you?"
for chunk in Chunker(text):
print(bytes(chunk))
# Custom size and delimiters
for chunk in Chunker(text, size=1024, delimiters=".?!\n"):
print(bytes(chunk))
Verify before relying
- Actual throughput benchmarks and whether '1 TB/s' claim is measured on specific hardware or representative of typical use.
- Memory overhead of the Chunker object and whether it scales linearly with text size.
- Behavior when text contains multi-byte UTF-8 sequences at chunk boundaries.
Package facts
| License | MIT OR Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 221 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 96,649 / month, #13,201 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: RustTopic :: Text Processing |
Evidence: memchunk-0.4.0-cp310-cp310-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; memchunk-0.4.0-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; memchunk-0.4.0-cp310-cp310-win_amd64.whl; memchunk-0.4.0-cp311-cp311-macosx_10_12_x86_64.whl; memchunk-0.4.0-cp311-cp311-macosx_11_0_arm64.whl; memchunk-0.4.0-cp311-cp311-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; memchunk-0.4.0-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; memchunk-0.4.0-cp311-cp311-win_amd64.whl; memchunk-0.4.0-cp312-cp312-macosx_10_12_x86_64.whl; memchunk-0.4.0-cp312-cp312-macosx_11_0_arm64.whl; memchunk-0.4.0-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; memchunk-0.4.0-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; memchunk-0.4.0-cp312-cp312-win_amd64.whl; memchunk-0.4.0-cp313-cp313-macosx_10_12_x86_64.whl; memchunk-0.4.0-cp313-cp313-macosx_11_0_arm64.whl; memchunk-0.4.0-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; memchunk-0.4.0-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; memchunk-0.4.0-cp313-cp313t-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; memchunk-0.4.0-cp313-cp313-win_amd64.whl; memchunk-0.4.0-cp314-cp314-macosx_10_12_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “text chunking library”
- memchunkSplits text into semantic chunks at delimiters (periods, newlines,…
- chonkieChonkie splits text into semantically meaningful chunks for RAG…
- chonkie-corechonkie-core splits text at semantic boundaries (periods, newlines,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also chonkie-core · semchunk · chonkie · semantic-text-splitter · langchain-text-splitters · sentence-stream · segtok · bytesparse · razdel · icechunk