tiktoken
tiktoken is a fast BPE tokeniser for use with OpenAI's models
Decision gist · record as of 2026-08-14
Yes. tiktoken is actively maintained, permissively licensed, has no known vulnerabilities, and is the canonical tokenizer for OpenAI models. Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models. Medium install friction is typical for compiled packages and poses no practical barrier on standard platforms.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later; compiled wheels available for cp310–cp312 on major platforms.
- Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp312, x86_64/aarch64/arm64, Linux/macOS/Windows).
- Actively maintained with recent commits and no known vulnerabilities.
License · maintenance · safety
permissive license (permissive) — MIT License permits commercial and private use, modification, and redistribution with minimal restrictions—suitable for most projects.
last release 2026-05-15 (91 days) · last repo commit 2026-05-24 · 18,987 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 233,042,766 downloads/mo, #172 on PyPI
Alternatives
Verify before relying
pip install tiktoken
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
tokens = enc.encode("hello world")
text = enc.decode(tokens)- Whether the educational submodule (tiktoken._educational) is production-ready or intended only for learning.
- Performance comparison claims (3-6x faster) are based on tiktoken==0.2.0; current performance relative to modern alternatives unknown.
What it is and what it does
tiktoken is OpenAI's official byte-pair encoding (BPE) tokenizer, designed to convert text into token sequences that language models consume. It provides fast, reversible encoding/decoding and supports multiple encoding schemes tied to specific OpenAI models (e.g., o200k_base for GPT-4o). The package depends on regex and requests for its core functionality.
The tokenizer is commonly used to count tokens before sending text to OpenAI APIs, to understand how models will segment input, and to work with custom encodings. It includes an educational submodule for learning BPE mechanics and a plugin mechanism (tiktoken_ext) for registering custom encodings.
Use it for
- Count tokens in text before calling OpenAI APIs to estimate costs and stay within context limits.
- Encode text into token sequences for analysis or debugging of how models will process input.
- Decode token IDs back to text for inspection or validation of model outputs.
- Build custom tokenizer encodings by extending tiktoken with domain-specific token vocabularies.
- Learn how byte-pair encoding works using the educational submodule and visualization tools.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
tiktoken is actively maintained, permissively licensed, has no known vulnerabilities, and is the canonical tokenizer for OpenAI models. Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models. Medium install friction is typical for compiled packages and poses no practical barrier on standard platforms.
Install
tiktoken on PyPI
Before you install
Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp312, x86_64/aarch64/arm64, Linux/macOS/Windows). Actively maintained with recent commits and no known vulnerabilities.
Requires Python 3.9 or later; compiled wheels available for cp310–cp312 on major platforms.
License in practice
MIT License permits commercial and private use, modification, and redistribution with minimal restrictions—suitable for most projects.
Quickstart
pip install tiktoken
import tiktoken
enc = tiktoken.get_encoding("o200k_base")
tokens = enc.encode("hello world")
text = enc.decode(tokens)
Verify before relying
- Whether the educational submodule (tiktoken._educational) is production-ready or intended only for learning.
- Performance comparison claims (3-6x faster) are based on tiktoken==0.2.0; current performance relative to modern alternatives unknown.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 2 packagesregexrequests |
| Maintenance | Actively maintained 91 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 233,042,766 / month, #172 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: tiktoken-0.13.0-cp310-cp310-macosx_10_12_x86_64.whl; tiktoken-0.13.0-cp310-cp310-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp310-cp310-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp310-cp310-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp310-cp310-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp310-cp310-musllinux_1_2_x86_64.whl; tiktoken-0.13.0-cp310-cp310-win_amd64.whl; tiktoken-0.13.0-cp311-cp311-macosx_10_12_x86_64.whl; tiktoken-0.13.0-cp311-cp311-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp311-cp311-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp311-cp311-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp311-cp311-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp311-cp311-musllinux_1_2_x86_64.whl; tiktoken-0.13.0-cp311-cp311-win_amd64.whl; tiktoken-0.13.0-cp312-cp312-macosx_10_13_x86_64.whl; tiktoken-0.13.0-cp312-cp312-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp312-cp312-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp312-cp312-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp312-cp312-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp312-cp312-musllinux_1_2_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “openai token counter”
- tiktokentiktoken is a fast BPE tokenizer that converts text into token…
- tokencostCounts tokens and estimates USD costs for LLM API calls across…
- openai-messages-token-helperEstimates token usage for OpenAI Chat Completions API messages and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
See also fastokens · tokie · tokenizers · pytorch-tokenizers · curated-tokenizers · openai-messages-token-helper · semchunk · blingfire · mikeshardmind-base2048 · tensorflow-text