triton
A language and compiler for custom Deep Learning operations
Decision gist · record as of 2026-08-14
Yes, if you need to write custom GPU kernels for deep learning. Triton significantly lowers the barrier to GPU programming compared to CUDA, with active maintenance, no security vulnerabilities, and permissive licensing. Install friction is moderate due to platform-specific wheels and optional build-from-source requirements, but pre-built wheels cover modern Python versions. Not necessary if you only use pre-built operations from frameworks like PyTorch.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a GPU (NVIDIA, AMD, or Intel) for execution; CPU-only systems can use the Triton interpreter (TRITON_INTERPRET=1) for development and testing.
- Medium install friction due to platform-specific wheels and compilation requirements.
- The package is actively maintained with recent releases and has a large community (19940 stars).
License · maintenance · safety
permissive license (permissive) — Triton is released under a permissive license, meaning you can use, modify, and distribute it with minimal restrictions in both open-source and commercial projects.
last release 2026-06-17 (58 days) · last repo commit 2026-08-14 · 19,940 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 72,441,274 downloads/mo, #462 on PyPI
Alternatives
Verify before relying
pip install triton
import triton
import triton.language as tl
@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(axis=0)
block_start = pid * BLOCK_SIZE
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask)
y = x * 2
tl.store(y_ptr + offsets, y, mask=mask)- Whether the package supports AMD and Intel GPUs equally well or if NVIDIA is the primary target.
- Performance benchmarks comparing Triton-compiled kernels to hand-written CUDA for common operations.
- Maturity level of the interpreter mode for development workflows without GPU access.
What it is and what it does
Triton is a compiler and language for writing custom deep-learning primitives that execute efficiently on GPUs. It sits between high-level frameworks like PyTorch and low-level CUDA, letting you write kernels in a Python-like syntax that compiles to optimized GPU code. The language abstracts away many low-level details (thread management, memory coalescing, tiling strategies) that CUDA requires you to handle manually, while still giving you fine-grained control over computation and memory access patterns.
Triton targets modern Python versions (3.10–3.14) and depends only on importlib-metadata at runtime. It is actively maintained, with a recent release cycle and strong community adoption. The package includes an interpreter mode for development and testing without a GPU, making it accessible for prototyping. Binary wheels are provided for common platforms, though building from source requires LLVM and additional build dependencies.
Use it for
- Implement custom CUDA kernels for neural network operations (attention, normalization, fused operations) without writing C++.
- Prototype and optimize GPU kernels for machine-learning inference and training pipelines.
- Develop platform-agnostic GPU code that can target NVIDIA, AMD, or Intel GPUs with minimal changes.
- Debug and iterate on GPU kernel logic using the Triton interpreter in development before deploying to hardware.
- Build domain-specific GPU operations for frameworks like PyTorch or JAX without leaving Python.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to write custom GPU kernels for deep learning.
Triton significantly lowers the barrier to GPU programming compared to CUDA, with active maintenance, no security vulnerabilities, and permissive licensing. Install friction is moderate due to platform-specific wheels and optional build-from-source requirements, but pre-built wheels cover modern Python versions. Not necessary if you only use pre-built operations from frameworks like PyTorch.
Install
triton on PyPI
Before you install
Medium install friction due to platform-specific wheels and compilation requirements. The package is actively maintained with recent releases and has a large community (19940 stars). Binary wheels are available for CPython 3.10–3.14 on x86_64 and aarch64, but building from source requires LLVM and build-time dependencies.
Requires a GPU (NVIDIA, AMD, or Intel) for execution; CPU-only systems can use the Triton interpreter (TRITON_INTERPRET=1) for development and testing.
License in practice
Triton is released under a permissive license, meaning you can use, modify, and distribute it with minimal restrictions in both open-source and commercial projects.
Quickstart
pip install triton
import triton
import triton.language as tl
@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
pid = tl.program_id(axis=0)
block_start = pid * BLOCK_SIZE
offsets = block_start + tl.arange(0, BLOCK_SIZE)
mask = offsets < n_elements
x = tl.load(x_ptr + offsets, mask=mask)
y = x * 2
tl.store(y_ptr + offsets, y, mask=mask)
Verify before relying
- Whether the package supports AMD and Intel GPUs equally well or if NVIDIA is the primary target.
- Performance benchmarks comparing Triton-compiled kernels to hand-written CUDA for common operations.
- Maturity level of the interpreter mode for development workflows without GPU access.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release <3.15,>=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 1 packageimportlib-metadata |
| Maintenance | Actively maintained 58 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 72,441,274 / month, #462 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Build Tools |
Evidence: triton-3.7.1-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “neural network primitives compiler”
- tritonTriton is a language and compiler for writing custom deep-learning…
- tokenspeed-tritonTriton is a language and compiler for writing custom deep-learning…
- nvidia-cudnn-cu12Provides cuDNN runtime libraries for GPU-accelerated deep neural…
Give your agent the search over MCP, or paste the wish link into any chat.
More Build Tools packages
Provides reusable utilities for Python packaging interoperability, including version handling, specifiers, markers, requirements, tags, and metadata parsing according to standards like PEP 440 and PEP 425.
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
pip is the standard installer for Python packages, enabling you to download and install packages from the Python Package Index and other indexes into your Python environment.
Hatchling is a standards-compliant Python build backend that handles packaging, metadata, and distribution of Python projects when configured in a project's pyproject.toml file.
Generates Python gRPC service stubs and message classes from Protocol Buffer definitions, enabling developers to build gRPC clients and servers.
pre-commit is a framework for installing and running git hooks written in any language before commits are made, automating code quality and validation checks across multi-language projects.
Install it if your team needs consistent, automated validation at commit time.
See also tokenspeed-triton · triton-windows · triton-ascend · tensorflow · nvidia-cudnn-cu12 · nvidia-cudnn-cu11 · nvidia-cudnn-cu13 · tf-nightly · perf-analyzer · pyston