$npx skillfedfor your agent

triton

A language and compiler for custom Deep Learning operations

With conditionsPyPI Build ToolsReleased Jun 202672.4M downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — triton-3.7.1-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl · triton-3.7.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl · triton-3.7.1-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl
v3.7.1 · released 2026-06-17 · Python <3.15,>=3.10 · 1 runtime deps: importlib-metadata

Yes, if you need to write custom GPU kernels for deep learning. Triton significantly lowers the barrier to GPU programming compared to CUDA, with active maintenance, no security vulnerabilities, and permissive licensing. Install friction is moderate due to platform-specific wheels and optional build-from-source requirements, but pre-built wheels cover modern Python versions. Not necessary if you only use pre-built operations from frameworks like PyTorch.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires a GPU (NVIDIA, AMD, or Intel) for execution; CPU-only systems can use the Triton interpreter (TRITON_INTERPRET=1) for development and testing.
  • Medium install friction due to platform-specific wheels and compilation requirements.
  • The package is actively maintained with recent releases and has a large community (19940 stars).

License · maintenance · safety

permissive license (permissive) — Triton is released under a permissive license, meaning you can use, modify, and distribute it with minimal restrictions in both open-source and commercial projects.

last release 2026-06-17 (58 days) · last repo commit 2026-08-14 · 19,940 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 72,441,274 downloads/mo, #462 on PyPI

Verify before relying

pip install triton

import triton
import triton.language as tl

@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(axis=0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    x = tl.load(x_ptr + offsets, mask=mask)
    y = x * 2
    tl.store(y_ptr + offsets, y, mask=mask)
  • Whether the package supports AMD and Intel GPUs equally well or if NVIDIA is the primary target.
  • Performance benchmarks comparing Triton-compiled kernels to hand-written CUDA for common operations.
  • Maturity level of the interpreter mode for development workflows without GPU access.
Same gist for agents: .md · .json

What it is and what it does

Triton is a compiler and language for writing custom deep-learning primitives that execute efficiently on GPUs. It sits between high-level frameworks like PyTorch and low-level CUDA, letting you write kernels in a Python-like syntax that compiles to optimized GPU code. The language abstracts away many low-level details (thread management, memory coalescing, tiling strategies) that CUDA requires you to handle manually, while still giving you fine-grained control over computation and memory access patterns.

Triton targets modern Python versions (3.10–3.14) and depends only on importlib-metadata at runtime. It is actively maintained, with a recent release cycle and strong community adoption. The package includes an interpreter mode for development and testing without a GPU, making it accessible for prototyping. Binary wheels are provided for common platforms, though building from source requires LLVM and additional build dependencies.

Use it for

  • Implement custom CUDA kernels for neural network operations (attention, normalization, fused operations) without writing C++.
  • Prototype and optimize GPU kernels for machine-learning inference and training pipelines.
  • Develop platform-agnostic GPU code that can target NVIDIA, AMD, or Intel GPUs with minimal changes.
  • Debug and iterate on GPU kernel logic using the Triton interpreter in development before deploying to hardware.
  • Build domain-specific GPU operations for frameworks like PyTorch or JAX without leaving Python.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to write custom GPU kernels for deep learning.

Triton significantly lowers the barrier to GPU programming compared to CUDA, with active maintenance, no security vulnerabilities, and permissive licensing. Install friction is moderate due to platform-specific wheels and optional build-from-source requirements, but pre-built wheels cover modern Python versions. Not necessary if you only use pre-built operations from frameworks like PyTorch.

Install

triton on PyPI

Before you install

Medium install friction due to platform-specific wheels and compilation requirements. The package is actively maintained with recent releases and has a large community (19940 stars). Binary wheels are available for CPython 3.10–3.14 on x86_64 and aarch64, but building from source requires LLVM and build-time dependencies.

Requires a GPU (NVIDIA, AMD, or Intel) for execution; CPU-only systems can use the Triton interpreter (TRITON_INTERPRET=1) for development and testing.

License in practice

Triton is released under a permissive license, meaning you can use, modify, and distribute it with minimal restrictions in both open-source and commercial projects.

Quickstart

pip install triton

import triton
import triton.language as tl

@triton.jit
def kernel(x_ptr, y_ptr, n_elements, BLOCK_SIZE: tl.constexpr):
    pid = tl.program_id(axis=0)
    block_start = pid * BLOCK_SIZE
    offsets = block_start + tl.arange(0, BLOCK_SIZE)
    mask = offsets < n_elements
    x = tl.load(x_ptr + offsets, mask=mask)
    y = x * 2
    tl.store(y_ptr + offsets, y, mask=mask)

Verify before relying

  • Whether the package supports AMD and Intel GPUs equally well or if NVIDIA is the primary target.
  • Performance benchmarks comparing Triton-compiled kernels to hand-written CUDA for common operations.
  • Maturity level of the interpreter mode for development workflows without GPU access.

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release <3.15,>=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependencies
1 package
importlib-metadata
MaintenanceActively maintained 58 days since the last release
Last repo commit
First released
Downloads72,441,274 / month, #462 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Software Development :: Build Tools

Evidence: triton-3.7.1-cp310-cp310-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp311-cp311-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp312-cp312-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp313-cp313-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp314-cp314-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl; triton-3.7.1-cp314-cp314t-manylinux_2_27_aarch64.manylinux_2_28_aarch64.whl; triton-3.7.1-cp314-cp314t-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Tags

Capabilities
GPU kernel compilercustom deep learning operationsCUDA alternativeneural network primitives compilerhigh-performance GPU programmingdeep learning kernel languagetiled computation compiler
Topics
gpu-compilerdeep-learningcuda-alternative
PyPI keywords
CompilerDeep Learning

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “neural network primitives compiler”

  • tritonTriton is a language and compiler for writing custom deep-learning…
  • tokenspeed-tritonTriton is a language and compiler for writing custom deep-learning…
  • nvidia-cudnn-cu12Provides cuDNN runtime libraries for GPU-accelerated deep neural…

Give your agent the search over MCP, or paste the wish link into any chat.

More Build Tools packages

packaging Worth it
PyPI · Build Tools · released Aug 2026

Provides reusable utilities for Python packaging interoperability, including version handling, specifiers, markers, requirements, tags, and metadata parsing according to standards like PEP 440 and PEP 425.

Apache-2.0 OR BSD-2-Clausepure Python · 3.9+
2.2Bdownloads / mo
tqdm Worth it
PyPI · Libraries · released Jul 2026

Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.

copyleftpure Python · 3.8+
648.6Mdownloads / mo
pip Worth it
PyPI · Build Tools · released Aug 2026

pip is the standard installer for Python packages, enabling you to download and install packages from the Python Package Index and other indexes into your Python environment.

MITpure Python · 3.10+
617.5Mdownloads / mo
hatchling Worth it
PyPI · Python Modules · released Aug 2026

Hatchling is a standards-compliant Python build backend that handles packaging, metadata, and distribution of Python projects when configured in a project's pyproject.toml file.

MITpure Python · 3.10+
484.2Mdownloads / mo
grpcio-tools Worth it
PyPI · Build Tools · released Jul 2026

Generates Python gRPC service stubs and message classes from Protocol Buffer definitions, enabling developers to build gRPC clients and servers.

Apache-2.0compiled wheel · 3.10+
278.2Mdownloads / mo
pre-commit Worth it
PyPI · Build Tools · released Aug 2026

pre-commit is a framework for installing and running git hooks written in any language before commits are made, automating code quality and validation checks across multi-language projects.

Install it if your team needs consistent, automated validation at commit time.

permissive licensepure Python · 3.10+
179.9Mdownloads / mo

See also tokenspeed-triton · triton-windows · triton-ascend · tensorflow · nvidia-cudnn-cu12 · nvidia-cudnn-cu11 · nvidia-cudnn-cu13 · tf-nightly · perf-analyzer · pyston

Further reading