chonkie
🦛 CHONK your texts with Chonkie ✨ - The no-nonsense chunking library
Decision gist · record as of 2026-08-14
Yes. Chonkie is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers low install friction with a focused set of dependencies. It solves a genuine problem in RAG pipelines—text chunking—with multiple strategies and integrations. The modular design lets you install only what you need, and the API server option adds deployment flexibility. Suitable for production RAG systems.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low friction install with a pure-Python wheel and six runtime dependencies.
- Active maintenance with recent commits and 4673 repository stars.
License · maintenance · safety
permissive license (permissive) — MIT License permits unrestricted use, modification, and redistribution in commercial and private projects with minimal obligations—only attribution and license inclusion required.
last release 2026-07-07 (38 days) · last repo commit 2026-08-08 · 4,673 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,346,996 downloads/mo, #4,020 on PyPI
Alternatives
Verify before relying
pip install chonkie
from chonkie import RecursiveChunker
chunker = RecursiveChunker()
chunks = chunker("Your text here")
for chunk in chunks:
print(chunk.text)- Whether all 45+ integrations mentioned in the description are available in the base install or require optional extras
- Performance characteristics of different chunkers (e.g., the claimed '100+ GB/s' for FastChunker)
- Multilingual support coverage across the stated 56 languages
What it is and what it does
Chonkie is a text chunking library designed to prepare documents for retrieval-augmented generation (RAG) systems. It provides multiple chunking strategies—including recursive, semantic, token-based, code-aware, and neural approaches—each suited to different content types and use cases. The library integrates with tokenizers, embedding models, vector databases, and LLMs, allowing you to build end-to-end pipelines that fetch, chunk, refine, embed, and load data into your RAG infrastructure.
The package emphasizes minimal dependencies by default: the base install includes only what's needed for common chunking tasks, with optional extras for specialized features like semantic chunking, code analysis, or API server deployment. It supports both synchronous and asynchronous workflows, can run as a self-hosted REST API, and handles text preprocessing through pluggable "Chef" components for markdown, tables, and OCR.
Use it for
- Split long documents into token-bounded chunks for LLM context windows before embedding and retrieval
- Chunk code repositories by syntactic structure to preserve function and class boundaries for code search
- Build a RAG pipeline that chains recursive chunking, semantic refinement, and embedding in a single workflow
- Run Chonkie as a microservice API to chunk documents from multiple applications without duplicating logic
- Process markdown or CSV files into structured chunks with context overlap for better retrieval quality
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Chonkie is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers low install friction with a focused set of dependencies. It solves a genuine problem in RAG pipelines—text chunking—with multiple strategies and integrations. The modular design lets you install only what you need, and the API server option adds deployment flexibility. Suitable for production RAG systems.
Install
chonkie on PyPI
Before you install
Low friction install with a pure-Python wheel and six runtime dependencies. Active maintenance with recent commits and 4673 repository stars. Supports Python 3.10 through 3.13.
Requires Python 3.10 or later.
License in practice
MIT License permits unrestricted use, modification, and redistribution in commercial and private projects with minimal obligations—only attribution and license inclusion required.
Quickstart
pip install chonkie
from chonkie import RecursiveChunker
chunker = RecursiveChunker()
chunks = chunker("Your text here")
for chunk in chunks:
print(chunk.text)
Verify before relying
- Whether all 45+ integrations mentioned in the description are available in the base install or require optional extras
- Performance characteristics of different chunkers (e.g., the claimed '100+ GB/s' for FastChunker)
- Multilingual support coverage across the stated 56 languages
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagestqdmnumpychonkie-coretenacityhttpxtokie |
| Maintenance | Actively maintained 38 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,346,996 / month, #4,020 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text Processing :: Linguistic |
Evidence: chonkie-1.7.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “token-based text segmentation”
- chonkieChonkie splits text into semantically meaningful chunks for RAG…
- pyobjc-framework-LatentSemanticMappingProvides Python bindings to macOS's LatentSemanticMapping framework…
- jieba3kPerforms Chinese word segmentation, breaking Chinese text into…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also chonkie-core · semchunk · memchunk · semantic-text-splitter · aurelio-sdk · langchain-text-splitters · langchain-graph-retriever · jieba3k · sentence-transformers · gitingest