$npx skillfedfor your agent

tokenspeed-triton

A language and compiler for custom Deep Learning operations (vendor release for TokenSpeed)

With conditionsPyPI Build ToolsReleased Jul 20262.0M downloads / moMITPlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl · tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl · tokenspeed_triton-3.8.10.post20260721-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl
v3.8.10.post20260721 · released 2026-07-21 · Python <3.15,>=3.10 · 1 runtime deps: importlib-metadata

Yes, if you are writing custom deep-learning operations and want to avoid low-level CUDA programming. The package is actively maintained, permissively licensed, has no known vulnerabilities, and runs on modern Python versions. Install friction is moderate due to binary wheels being available but source builds requiring LLVM; this is a one-time cost. Not recommended if you only use standard framework operations or lack GPU hardware for testing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a compatible GPU (NVIDIA or AMD) and CUDA/ROCm runtime for execution; CPU-only testing is possible with TRITON_INTERPRET=1 environment variable.
  • Medium install friction: wheels are available for Python 3.10–3.14 on x86_64 and aarch64 Linux, but building from source requires LLVM and build-time dependencies.
  • The package is actively maintained with a recent release (24 days old) and a well-established repository.

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.

last release 2026-07-21 (24 days) · last repo commit 2026-08-14 · 19,944 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,990,943 downloads/mo, #3,380 on PyPI

Verify before relying

pip install tokenspeed-triton

import triton
import triton.language as tl

@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(axis=0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    x = tl.load(x_ptr + offsets, mask=mask)
    y = x * 2
    tl.store(y_ptr + offsets, y, mask=mask)
  • Whether prebuilt wheels cover all target deployment architectures and CUDA versions needed for your use case.
  • Performance gains relative to hand-written CUDA or other kernel frameworks for your specific workloads.
  • Maturity of AMD GPU support and whether it matches NVIDIA feature parity.
Same gist for agents: .md · .json

What it is and what it does

Triton is a compiler and language designed to let you write high-performance deep-learning kernels at a higher level of abstraction than CUDA, while retaining the ability to optimize for specific hardware. Instead of writing low-level GPU code, you write Triton kernels in Python-like syntax, and the compiler handles the translation to efficient machine code for GPUs and CPUs. It sits between the productivity of high-level frameworks and the control of hand-written CUDA.

The package is intended for researchers and engineers building custom neural-network operations—operations that existing frameworks don't provide or that need domain-specific optimization. It includes a just-in-time compiler, an interpreter for CPU-based testing, and support for tiled computation patterns common in deep learning. The main runtime dependency is importlib-metadata; the package itself handles LLVM integration during build time.

Use it for

  • Write custom CUDA-like kernels for novel neural-network layers without learning low-level GPU programming.
  • Optimize matrix operations, attention mechanisms, or other primitives for specific hardware without rewriting in C++.
  • Prototype and test GPU kernels on CPU using the Triton interpreter before deploying to hardware.
  • Build domain-specific deep-learning libraries that need fine-grained control over memory and compute.
  • Accelerate research by rapidly iterating on custom operations in a higher-level language than CUDA.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are writing custom deep-learning operations and want to avoid low-level CUDA programming.

The package is actively maintained, permissively licensed, has no known vulnerabilities, and runs on modern Python versions. Install friction is moderate due to binary wheels being available but source builds requiring LLVM; this is a one-time cost. Not recommended if you only use standard framework operations or lack GPU hardware for testing.

Install

tokenspeed-triton on PyPI

Before you install

Medium install friction: wheels are available for Python 3.10–3.14 on x86_64 and aarch64 Linux, but building from source requires LLVM and build-time dependencies. The package is actively maintained with a recent release (24 days old) and a well-established repository.

Requires a compatible GPU (NVIDIA or AMD) and CUDA/ROCm runtime for execution; CPU-only testing is possible with TRITON_INTERPRET=1 environment variable.

License in practice

MIT license is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.

Quickstart

pip install tokenspeed-triton

import triton
import triton.language as tl

@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(axis=0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    x = tl.load(x_ptr + offsets, mask=mask)
    y = x * 2
    tl.store(y_ptr + offsets, y, mask=mask)

Verify before relying

  • Whether prebuilt wheels cover all target deployment architectures and CUDA versions needed for your use case.
  • Performance gains relative to hand-written CUDA or other kernel frameworks for your specific workloads.
  • Maturity of AMD GPU support and whether it matches NVIDIA feature parity.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <3.15,>=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
1 package
importlib-metadata
MaintenanceActively maintained 24 days since the last release
Last repo commit
First released
Downloads1,990,943 / month, #3,380 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Build Tools

Evidence: tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; tokenspeed_triton-3.8.10.post20260721-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; tokenspeed_triton-3.8.10.post20260721-cp312-abi3-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; tokenspeed_triton-3.8.10.post20260721-cp312-abi3-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Tags

Capabilities
GPU kernel compiler for deep learningcustom neural network operationsCUDA alternative high-leveldeep learning kernel languageefficient GPU code generationML primitive compilertiled neural network computations
Topics
gpu-kernel-compilerdeep-learning-infrastructure
PyPI keywords
CompilerDeep Learning

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “custom neural network operations”

  • tokenspeed-tritonTriton is a language and compiler for writing custom deep-learning…
  • cuequivariance-ops-torch-cu12Provides CUDA-accelerated PyTorch kernels and operators for…
  • thopCounts floating-point operations (MACs) and parameters in PyTorch…

Give your agent the search over MCP, or paste the wish link into any chat.

More Build Tools packages

packaging Worth it
PyPI · Build Tools · released Aug 2026

Provides reusable utilities for Python packaging interoperability, including version handling, specifiers, markers, requirements, tags, and metadata parsing according to standards like PEP 440 and PEP 425.

Apache-2.0 OR BSD-2-Clausepure Python · 3.9+
2.2Bdownloads / mo
tqdm Worth it
PyPI · Libraries · released Jul 2026

Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.

copyleftpure Python · 3.8+
648.6Mdownloads / mo
pip Worth it
PyPI · Build Tools · released Aug 2026

pip is the standard installer for Python packages, enabling you to download and install packages from the Python Package Index and other indexes into your Python environment.

MITpure Python · 3.10+
617.5Mdownloads / mo
hatchling Worth it
PyPI · Python Modules · released Aug 2026

Hatchling is a standards-compliant Python build backend that handles packaging, metadata, and distribution of Python projects when configured in a project's pyproject.toml file.

MITpure Python · 3.10+
484.2Mdownloads / mo
grpcio-tools Worth it
PyPI · Build Tools · released Jul 2026

Generates Python gRPC service stubs and message classes from Protocol Buffer definitions, enabling developers to build gRPC clients and servers.

Apache-2.0compiled wheel · 3.10+
278.2Mdownloads / mo
pre-commit Worth it
PyPI · Build Tools · released Aug 2026

pre-commit is a framework for installing and running git hooks written in any language before commits are made, automating code quality and validation checks across multi-language projects.

Install it if your team needs consistent, automated validation at commit time.

permissive licensepure Python · 3.10+
179.9Mdownloads / mo

See also triton · triton-windows · triton-ascend · numba · helion · liger-kernel · perf-analyzer · flash-attn · nvidia-nvvm · pyston

Further reading