$npx skillfedfor your agent

tiktoken

tiktoken is a fast BPE tokeniser for use with OpenAI's models

Worth itPyPI LinguisticReleased May 2026233.0M downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — tiktoken-0.13.0-cp310-cp310-macosx_10_12_x86_64.whl · tiktoken-0.13.0-cp310-cp310-macosx_11_0_arm64.whl · tiktoken-0.13.0-cp310-cp310-manylinux_2_28_aarch64.whl
v0.13.0 · released 2026-05-15 · Python >=3.9 · 2 runtime deps: regex, requests

Yes. tiktoken is actively maintained, permissively licensed, has no known vulnerabilities, and is the canonical tokenizer for OpenAI models. Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models. Medium install friction is typical for compiled packages and poses no practical barrier on standard platforms.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later; compiled wheels available for cp310–cp312 on major platforms.
  • Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp312, x86_64/aarch64/arm64, Linux/macOS/Windows).
  • Actively maintained with recent commits and no known vulnerabilities.

License · maintenance · safety

permissive license (permissive) — MIT License permits commercial and private use, modification, and redistribution with minimal restrictions—suitable for most projects.

last release 2026-05-15 (91 days) · last repo commit 2026-05-24 · 18,987 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 233,042,766 downloads/mo, #172 on PyPI

Verify before relying

pip install tiktoken

import tiktoken
enc = tiktoken.get_encoding("o200k_base")
tokens = enc.encode("hello world")
text = enc.decode(tokens)
  • Whether the educational submodule (tiktoken._educational) is production-ready or intended only for learning.
  • Performance comparison claims (3-6x faster) are based on tiktoken==0.2.0; current performance relative to modern alternatives unknown.
Same gist for agents: .md · .json

What it is and what it does

tiktoken is OpenAI's official byte-pair encoding (BPE) tokenizer, designed to convert text into token sequences that language models consume. It provides fast, reversible encoding/decoding and supports multiple encoding schemes tied to specific OpenAI models (e.g., o200k_base for GPT-4o). The package depends on regex and requests for its core functionality.

The tokenizer is commonly used to count tokens before sending text to OpenAI APIs, to understand how models will segment input, and to work with custom encodings. It includes an educational submodule for learning BPE mechanics and a plugin mechanism (tiktoken_ext) for registering custom encodings.

Use it for

  • Count tokens in text before calling OpenAI APIs to estimate costs and stay within context limits.
  • Encode text into token sequences for analysis or debugging of how models will process input.
  • Decode token IDs back to text for inspection or validation of model outputs.
  • Build custom tokenizer encodings by extending tiktoken with domain-specific token vocabularies.
  • Learn how byte-pair encoding works using the educational submodule and visualization tools.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

tiktoken is actively maintained, permissively licensed, has no known vulnerabilities, and is the canonical tokenizer for OpenAI models. Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models. Medium install friction is typical for compiled packages and poses no practical barrier on standard platforms.

Install

tiktoken on PyPI

Before you install

Medium install friction due to compiled wheels for multiple Python versions and platforms (cp310–cp312, x86_64/aarch64/arm64, Linux/macOS/Windows). Actively maintained with recent commits and no known vulnerabilities.

Requires Python 3.9 or later; compiled wheels available for cp310–cp312 on major platforms.

License in practice

MIT License permits commercial and private use, modification, and redistribution with minimal restrictions—suitable for most projects.

Quickstart

pip install tiktoken

import tiktoken
enc = tiktoken.get_encoding("o200k_base")
tokens = enc.encode("hello world")
text = enc.decode(tokens)

Verify before relying

  • Whether the educational submodule (tiktoken._educational) is production-ready or intended only for learning.
  • Performance comparison claims (3-6x faster) are based on tiktoken==0.2.0; current performance relative to modern alternatives unknown.

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.9
Install frictionMedium. Platform-specific wheel
Runtime dependencies
2 packages
regexrequests
MaintenanceActively maintained 91 days since the last release
Last repo commit
First released
Downloads233,042,766 / month, #172 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: tiktoken-0.13.0-cp310-cp310-macosx_10_12_x86_64.whl; tiktoken-0.13.0-cp310-cp310-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp310-cp310-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp310-cp310-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp310-cp310-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp310-cp310-musllinux_1_2_x86_64.whl; tiktoken-0.13.0-cp310-cp310-win_amd64.whl; tiktoken-0.13.0-cp311-cp311-macosx_10_12_x86_64.whl; tiktoken-0.13.0-cp311-cp311-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp311-cp311-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp311-cp311-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp311-cp311-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp311-cp311-musllinux_1_2_x86_64.whl; tiktoken-0.13.0-cp311-cp311-win_amd64.whl; tiktoken-0.13.0-cp312-cp312-macosx_10_13_x86_64.whl; tiktoken-0.13.0-cp312-cp312-macosx_11_0_arm64.whl; tiktoken-0.13.0-cp312-cp312-manylinux_2_28_aarch64.whl; tiktoken-0.13.0-cp312-cp312-manylinux_2_28_x86_64.whl; tiktoken-0.13.0-cp312-cp312-musllinux_1_2_aarch64.whl; tiktoken-0.13.0-cp312-cp312-musllinux_1_2_x86_64.whl

Tags

Capabilities
openai token counterbpe tokenizertext to tokens conversiongpt tokenizationtoken encoding for language models
Topics
openai-integrationnlp-tokenizationbpe-encoding

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “openai token counter”

  • tiktokentiktoken is a fast BPE tokenizer that converts text into token…
  • tokencostCounts tokens and estimates USD costs for LLM API calls across…
  • openai-messages-token-helperEstimates token usage for OpenAI Chat Completions API messages and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Linguistic packages

charset-normalizer Worth it
PyPI · Utilities · released Aug 2026

Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.

permissive licensepure Python · 3.7+
1.7Bdownloads / mo
chardet Worth it
PyPI · Python Modules · released Aug 2026

Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.

Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.

0BSDpure Python · 3.10+
199.0Mdownloads / mo
text-unidecode With conditions
PyPI · Python Modules · released Aug 2019

Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.

However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.

GPL-2.0-or-laterpure Pythonabandoned
89.0Mdownloads / mo
lark Worth it
PyPI · Python Modules · released Oct 2025

Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.

MITpure Python · 3.8+
79.7Mdownloads / mo
tree-sitter Worth it
PyPI · Linguistic · released Jun 2026

Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.

MITcompiled wheel · 3.10+
79.0Mdownloads / mo
nltk Worth it
PyPI · Scientific/Engineering · released Aug 2026

NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.

Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…

Apache-2.0pure Python · 3.10+
74.1Mdownloads / mo

See also fastokens · tokie · tokenizers · pytorch-tokenizers · curated-tokenizers · openai-messages-token-helper · semchunk · blingfire · mikeshardmind-base2048 · tensorflow-text