pytokens
A Fast, spec compliant Python 3.14+ tokenizer that runs on older Pythons.
Decision gist · record as of 2026-08-14
Yes, if you need a spec-compliant Python tokenizer. The package is actively maintained, has no dependencies, carries a permissive MIT license, and ranks in the top 1000 PyPI packages by download volume. Install it if your tool requires tokenization; skip it if you only need AST-level analysis (use the standard library ast module instead).AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with no runtime dependencies.
- Actively maintained as of 2026-02-18 with recent release activity.
- Distributed as a wheel, optionally compiled with mypyc for performance.
License · maintenance · safety
permissive license (permissive) — MIT License permits unrestricted use, modification, and distribution with only attribution and liability waiver required—no restrictions on commercial or proprietary use.
last release 2026-01-30 (196 days) · last repo commit 2026-02-18 · 4 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 108,727,382 downloads/mo, #323 on PyPI
Alternatives
Verify before relying
pip install pytokens
python -m pytokens path/to/file.py- Whether the tokenizer output format and API are documented for programmatic use beyond the CLI
- Performance characteristics and whether mypyc compilation provides measurable speedup for typical workloads
- Whether the package is used as a dependency by other tools or primarily as a standalone utility
What it is and what it does
pytokens is a tokenizer for Python source code that implements the Python language specification. It can run as a compiled module on modern Python versions or fall back to pure Python on older interpreters, supporting Python 3.8 through 3.14. The package is designed to parse Python files into a token stream, which is useful for tools that need to analyze or manipulate Python source code at the lexical level.
The primary interface is a command-line tool that reads a Python file and outputs its tokens. The package optionally compiles with mypyc for performance, though this can be disabled via environment variable. It has no external runtime dependencies, making it straightforward to integrate into other projects that need tokenization capabilities.
Use it for
- Build linters or code analysis tools that need to examine Python source at the token level
- Create code formatters or pretty-printers that parse and regenerate Python source
- Develop IDE features like syntax highlighting or bracket matching that require token-level parsing
- Implement custom Python syntax validators or style checkers
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need a spec-compliant Python tokenizer.
The package is actively maintained, has no dependencies, carries a permissive MIT license, and ranks in the top 1000 PyPI packages by download volume. Install it if your tool requires tokenization; skip it if you only need AST-level analysis (use the standard library ast module instead).
Install
pytokens on PyPI
Before you install
Low install friction with no runtime dependencies. Actively maintained as of 2026-02-18 with recent release activity. Distributed as a wheel, optionally compiled with mypyc for performance.
License in practice
MIT License permits unrestricted use, modification, and distribution with only attribution and liability waiver required—no restrictions on commercial or proprietary use.
Quickstart
pip install pytokens
python -m pytokens path/to/file.py
Verify before relying
- Whether the tokenizer output format and API are documented for programmatic use beyond the CLI
- Performance characteristics and whether mypyc compilation provides measurable speedup for typical workloads
- Whether the package is used as a dependency by other tools or primarily as a standalone utility
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 196 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 108,727,382 / month, #323 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonTyping :: Typed |
Evidence: pytokens-0.4.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “parse python source code”
- pytokenspytokens is a Python tokenizer that parses Python source code into…
- ast-commentsExtends Python's built-in `ast` module to preserve comments as nodes…
- javalangjavalang is a pure Python lexer and parser for Java 8 source code…
Give your agent the search over MCP, or paste the wish link into any chat.
Similar packages
javalang is a pure Python lexer and parser for Java 8 source code that builds an abstract syntax tree you can traverse to extract information about Java classes, methods, and other language elements.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
fastokens is a high-performance BPE tokenizer for large language models, built on a Rust backend and compatible with HuggingFace tokenizer.json and tiktoken model formats.
However, verify the license status in the repository before use in proprietary projects, and confirm that unsupported tokenizer features do not block your use case.
A Rust-based JSON tokenizer that accelerates JSON parsing for the json-stream library, available as a drop-in replacement or automatic dependency.
PLY is a lexer and parser generator for Python that implements lex and yacc functionality using LALR(1) parsing, enabling you to build language processors and domain-specific language parsers entirely in Python.
However, do not use it for new production systems: the abandonment since 2018-02-15 means no modern Python compatibility assurance, no security updates, and no modern…
Parses Python source files and serializes their AST to mypy's native binary format using a Rust-backed parser, replacing the stdlib ast module for mypy's internal use.
See also tensorflow-text · tokenizer · curated-tokenizers · tokie