$npx skillfedfor your agent

fastokens

With conditionsPyPI LinguisticReleased Aug 2026551.2K downloads / moPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — fastokens-0.3.1-cp39-abi3-macosx_10_12_x86_64.whl · fastokens-0.3.1-cp39-abi3-macosx_11_0_arm64.whl · fastokens-0.3.1-cp39-abi3-manylinux_2_17_i686.manylinux2014_i686.whl
v0.3.1 · released 2026-08-04 · Python >=3.9

Yes, if you need faster tokenization in LLM inference and your model is among the tested families (Qwen, Kimi, Minimax, Gemma, DeepSeek) or you verify compatibility. The prebuilt wheels and zero runtime dependencies make installation frictionless. However, verify the license status in the repository before use in proprietary projects, and confirm that unsupported tokenizer features do not block your use case.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later.
  • Prebuilt wheels for Python 3.9+ across Linux, macOS, and Windows minimize installation friction.
  • The package is actively maintained with a recent release and no known vulnerabilities.

License · maintenance · safety

(unclear) — License status is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the repository's LICENSE file before adopting in proprietary or copyleft-sensitive projects.

last release 2026-08-04 (10 days) · last repo commit 2026-08-10 · 132 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 551,217 downloads/mo, #6,051 on PyPI

Verify before relying

pip install fastokens

from fastokens._native import Tokenizer
tokenizer = Tokenizer.from_model("deepseek-ai/DeepSeek-V3.2")
tokens = tokenizer.encode("A very long prompt that is now lightning fast.")
  • Whether the 10x+ performance improvement claim applies to your specific workload and model size.
  • Exact scope of features not supported compared to the tokenizers library.
  • Whether your target model is among the tested families (Qwen, Kimi, Minimax, Gemma, DeepSeek) or requires verification.
Same gist for agents: .md · .json

What it is and what it does

fastokens is a Rust-backed tokenizer designed to accelerate BPE encoding for large language models. It loads both HuggingFace tokenizer.json files and tiktoken model files, and ships prebuilt wheels for Python 3.9+ across major platforms, eliminating the need to compile Rust code during installation.

The main use case is speeding up tokenization in inference pipelines where prompt encoding becomes a bottleneck. It includes features like vocabulary extension through add_tokens and add_special_tokens, optional prefix caching for shared system prompts, and PCRE2 resource limits to guard against pathological regex patterns. The package can be used standalone or integrated with serving frameworks.

Use it for

  • Accelerate tokenization in inference pipelines to reduce encoding latency on large prompts.
  • Tokenize long contexts faster by enabling the optional prefix cache for repeated system prompts.
  • Load and use tiktoken models (e.g., cl100k_base or o200k_base) directly without additional dependencies.
  • Extend model vocabularies by adding placeholder tokens to match padded embedding matrices.
  • Replace default tokenizers in production serving to improve time-to-first-token metrics.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need faster tokenization in LLM inference and your model is among the tested families (Qwen, Kimi, Minimax, Gemma, DeepSeek) or you verify compatibility.

The prebuilt wheels and zero runtime dependencies make installation frictionless. However, verify the license status in the repository before use in proprietary projects, and confirm that unsupported tokenizer features do not block your use case.

Install

fastokens on PyPI

Before you install

Prebuilt wheels for Python 3.9+ across Linux, macOS, and Windows minimize installation friction. The package is actively maintained with a recent release and no known vulnerabilities.

Requires Python 3.9 or later.

License in practice

License status is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the repository's LICENSE file before adopting in proprietary or copyleft-sensitive projects.

Quickstart

pip install fastokens

from fastokens._native import Tokenizer
tokenizer = Tokenizer.from_model("deepseek-ai/DeepSeek-V3.2")
tokens = tokenizer.encode("A very long prompt that is now lightning fast.")

Verify before relying

  • Whether the 10x+ performance improvement claim applies to your specific workload and model size.
  • Exact scope of features not supported compared to the tokenizers library.
  • Whether your target model is among the tested families (Qwen, Kimi, Minimax, Gemma, DeepSeek) or requires verification.

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 10 days since the last release
Last repo commit
First released
Downloads551,217 / month, #6,051 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: fastokens-0.3.1-cp39-abi3-macosx_10_12_x86_64.whl; fastokens-0.3.1-cp39-abi3-macosx_11_0_arm64.whl; fastokens-0.3.1-cp39-abi3-manylinux_2_17_i686.manylinux2014_i686.whl; fastokens-0.3.1-cp39-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl; fastokens-0.3.1-cp39-abi3-manylinux_2_28_aarch64.whl; fastokens-0.3.1-cp39-abi3-manylinux_2_28_armv7l.whl; fastokens-0.3.1-cp39-abi3-manylinux_2_28_ppc64le.whl; fastokens-0.3.1-cp39-abi3-musllinux_1_2_aarch64.whl; fastokens-0.3.1-cp39-abi3-musllinux_1_2_armv7l.whl; fastokens-0.3.1-cp39-abi3-musllinux_1_2_i686.whl; fastokens-0.3.1-cp39-abi3-musllinux_1_2_x86_64.whl; fastokens-0.3.1-cp39-abi3-win_amd64.whl

Tags

Capabilities
fast BPE tokenizerLLM tokenizationtiktoken alternativehigh-speed text encodingbyte pair encoding
Topics
llm-inferencetokenizationperformance

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “tiktoken alternative”

  • fastokensfastokens is a high-performance BPE tokenizer for large language…
  • pytorch-tokenizersProvides C++ implementations of multiple tokenizers (SentencePiece,…
  • tiktokentiktoken is a fast BPE tokenizer that converts text into token…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
tiktoken Worth it
PyPI · Linguistic · released May 2026

tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.

Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.

permissive licensecompiled wheel · 3.9+
233.0Mdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo

See also tokie · tokenizers · pytorch-tokenizers · json-stream-rs-tokenizer · blingfire · pytokens · curated-tokenizers · toons · outlines-core