tokenspeed-triton
A language and compiler for custom Deep Learning operations (vendor release for TokenSpeed)
Decision gist · record as of 2026-08-14
Yes, if you are writing custom deep-learning operations and want to avoid low-level CUDA programming. The package is actively maintained, permissively licensed, has no known vulnerabilities, and runs on modern Python versions. Install friction is moderate due to binary wheels being available but source builds requiring LLVM; this is a one-time cost. Not recommended if you only use standard framework operations or lack GPU hardware for testing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a compatible GPU (NVIDIA or AMD) and CUDA/ROCm runtime for execution; CPU-only testing is possible with TRITON_INTERPRET=1 environment variable.
- Medium install friction: wheels are available for Python 3.10–3.14 on x86_64 and aarch64 Linux, but building from source requires LLVM and build-time dependencies.
- The package is actively maintained with a recent release (24 days old) and a well-established repository.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.
last release 2026-07-21 (24 days) · last repo commit 2026-08-14 · 19,944 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,990,943 downloads/mo, #3,380 on PyPI
Alternatives
Verify before relying
pip install tokenspeed-triton
import triton
import triton.language as tl
@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(axis=0)
block_start = pid * BLOCK_SIZE
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask)
y = x * 2
tl.store(y_ptr + offsets, y, mask=mask)- Whether prebuilt wheels cover all target deployment architectures and CUDA versions needed for your use case.
- Performance gains relative to hand-written CUDA or other kernel frameworks for your specific workloads.
- Maturity of AMD GPU support and whether it matches NVIDIA feature parity.
What it is and what it does
Triton is a compiler and language designed to let you write high-performance deep-learning kernels at a higher level of abstraction than CUDA, while retaining the ability to optimize for specific hardware. Instead of writing low-level GPU code, you write Triton kernels in Python-like syntax, and the compiler handles the translation to efficient machine code for GPUs and CPUs. It sits between the productivity of high-level frameworks and the control of hand-written CUDA.
The package is intended for researchers and engineers building custom neural-network operations—operations that existing frameworks don't provide or that need domain-specific optimization. It includes a just-in-time compiler, an interpreter for CPU-based testing, and support for tiled computation patterns common in deep learning. The main runtime dependency is importlib-metadata; the package itself handles LLVM integration during build time.
Use it for
- Write custom CUDA-like kernels for novel neural-network layers without learning low-level GPU programming.
- Optimize matrix operations, attention mechanisms, or other primitives for specific hardware without rewriting in C++.
- Prototype and test GPU kernels on CPU using the Triton interpreter before deploying to hardware.
- Build domain-specific deep-learning libraries that need fine-grained control over memory and compute.
- Accelerate research by rapidly iterating on custom operations in a higher-level language than CUDA.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are writing custom deep-learning operations and want to avoid low-level CUDA programming.
The package is actively maintained, permissively licensed, has no known vulnerabilities, and runs on modern Python versions. Install friction is moderate due to binary wheels being available but source builds requiring LLVM; this is a one-time cost. Not recommended if you only use standard framework operations or lack GPU hardware for testing.
Install
tokenspeed-triton on PyPI
Before you install
Medium install friction: wheels are available for Python 3.10–3.14 on x86_64 and aarch64 Linux, but building from source requires LLVM and build-time dependencies. The package is actively maintained with a recent release (24 days old) and a well-established repository.
Requires a compatible GPU (NVIDIA or AMD) and CUDA/ROCm runtime for execution; CPU-only testing is possible with TRITON_INTERPRET=1 environment variable.
License in practice
MIT license is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.
Quickstart
pip install tokenspeed-triton
import triton
import triton.language as tl
@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(axis=0)
block_start = pid * BLOCK_SIZE
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask)
y = x * 2
tl.store(y_ptr + offsets, y, mask=mask)
Verify before relying
- Whether prebuilt wheels cover all target deployment architectures and CUDA versions needed for your use case.
- Performance gains relative to hand-written CUDA or other kernel frameworks for your specific workloads.
- Maturity of AMD GPU support and whether it matches NVIDIA feature parity.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release <3.15,>=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packageimportlib-metadata |
| Maintenance | Actively maintained 24 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,990,943 / month, #3,380 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Build Tools |
Evidence: tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; tokenspeed_triton-3.8.10.post20260721-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; tokenspeed_triton-3.8.10.post20260721-cp312-abi3-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp312-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “custom neural network operations”
- tokenspeed-tritonTriton is a language and compiler for writing custom deep-learning…
- cuequivariance-ops-torch-cu12Provides CUDA-accelerated PyTorch kernels and operators for…
- thopCounts floating-point operations (MACs) and parameters in PyTorch…
Give your agent the search over MCP, or paste the wish link into any chat.
More Build Tools packages
Provides reusable utilities for Python packaging interoperability, including version handling, specifiers, markers, requirements, tags, and metadata parsing according to standards like PEP 440 and PEP 425.
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
pip is the standard installer for Python packages, enabling you to download and install packages from the Python Package Index and other indexes into your Python environment.
Hatchling is a standards-compliant Python build backend that handles packaging, metadata, and distribution of Python projects when configured in a project's pyproject.toml file.
Generates Python gRPC service stubs and message classes from Protocol Buffer definitions, enabling developers to build gRPC clients and servers.
pre-commit is a framework for installing and running git hooks written in any language before commits are made, automating code quality and validation checks across multi-language projects.
Install it if your team needs consistent, automated validation at commit time.
See also triton · triton-windows · triton-ascend · numba · helion · liger-kernel · perf-analyzer · flash-attn · nvidia-nvvm · pyston