helion
A Python-embedded DSL that makes it easy to write ML kernels
Decision gist · record as of 2026-08-14
Yes, if you need to write custom GPU kernels and want to avoid manual tuning. Low install friction, active maintenance, and zero known vulnerabilities support adoption. However, the unclear license classification and 10-minute autotuning overhead on first run are real constraints—verify license compatibility for your use case and expect startup latency. Best suited for teams with GPU access and kernels that justify the autotuning investment.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10 and a CUDA-capable GPU; first kernel execution triggers autotuning which takes approximately 10 minutes.
- Low install friction with a pure-Python wheel.
- Active maintenance with recent releases; the project shows active development as of 2026-08-14 with 922 repository stars.
License · maintenance · safety
(unclear) — License treatment is unclear; the package carries a BSD-style license from Meta Platforms but SPDX classification is not provided. Review the license text before use in proprietary or commercial contexts.
last release 2026-07-29 (16 days) · last repo commit 2026-08-14 · 922 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 204,427 downloads/mo, #9,607 on PyPI
Alternatives
Verify before relying
import helion
import helion.language as hl
@helion.kernel()
def matmul(x, y):
m, k = x.size()
k, n = y.size()
out = torch.empty([m, n], dtype=x.dtype, device=x.device)
for tile_m, tile_n in hl.tile([m, n]):
acc = hl.zeros([tile_m, tile_n], dtype=torch.float32)
for tile_k in hl.tile(k):
acc = torch.addmm(acc, x[tile_m, tile_k], y[tile_k, tile_n])
out[tile_m, tile_n] = acc
return out- Whether autotuning results are cached across runs and how to manage the cache.
- Supported operations and coverage limits beyond the documented examples.
- Performance overhead of the Helion compilation and autotuning pipeline.
- Compatibility with non-NVIDIA GPUs despite the Triton backend.
What it is and what it does
Helion is a higher-level abstraction over Triton that lets you write GPU kernels using familiar syntax, then automatically optimizes them through an extensive search process. Instead of manually tuning tile sizes, grid dimensions, memory access patterns, and kernel configurations, you write a kernel using operations inside Helion's tiling loops, and the system generates and evaluates hundreds of candidate implementations to find the fastest one for your hardware.
The package compiles code inside `@helion.kernel()` decorated functions into a single optimized GPU kernel. It automates decisions about tensor indexing strategies, masking, grid layout, loop reordering, warp specialization, and persistent kernel strategies. First execution triggers autotuning (typically around 10 minutes), after which you can hardcode the best configuration to skip re-tuning on subsequent runs.
Use it for
- Write custom matrix multiplication kernels without manually tuning configurations for each GPU architecture.
- Optimize reduction operations by letting Helion automatically choose loop strategies and memory access patterns.
- Develop portable GPU kernels that perform well across different hardware through broad search space exploration.
- Prototype GPU-accelerated operations before committing to hand-tuned implementations.
- Automate kernel argument handling and closure lifting for complex tensor operations.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to write custom GPU kernels and want to avoid manual tuning.
Low install friction, active maintenance, and zero known vulnerabilities support adoption. However, the unclear license classification and 10-minute autotuning overhead on first run are real constraints—verify license compatibility for your use case and expect startup latency. Best suited for teams with GPU access and kernels that justify the autotuning investment.
Install
helion on PyPI
Before you install
Low install friction with a pure-Python wheel. Active maintenance with recent releases; the project shows active development as of 2026-08-14 with 922 repository stars.
Requires Python >=3.10 and a CUDA-capable GPU; first kernel execution triggers autotuning which takes approximately 10 minutes.
License in practice
License treatment is unclear; the package carries a BSD-style license from Meta Platforms but SPDX classification is not provided. Review the license text before use in proprietary or commercial contexts.
Quickstart
import helion
import helion.language as hl
@helion.kernel()
def matmul(x, y):
m, k = x.size()
k, n = y.size()
out = torch.empty([m, n], dtype=x.dtype, device=x.device)
for tile_m, tile_n in hl.tile([m, n]):
acc = hl.zeros([tile_m, tile_n], dtype=torch.float32)
for tile_k in hl.tile(k):
acc = torch.addmm(acc, x[tile_m, tile_k], y[tile_k, tile_n])
out[tile_m, tile_n] = acc
return out
Verify before relying
- Whether autotuning results are cached across runs and how to manage the cache.
- Supported operations and coverage limits beyond the documented examples.
- Performance overhead of the Helion compilation and autotuning pipeline.
- Compatibility with non-NVIDIA GPUs despite the Triton backend.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesfilechecknumpypsutilrichscikit-learntyping-extensions |
| Maintenance | Actively maintained 16 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 204,427 / month, #9,607 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Operating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: helion-1.4.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “machine learning kernel compiler”
- helionHelion is a Python-embedded domain-specific language for writing…
- nvidia-cuda-nvccProvides the NVIDIA CUDA nvcc compiler for building CUDA applications…
- triton-ascendTriton Ascend is a compiler framework that enables Triton code to run…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also triton-ascend · liger-kernel · triton-windows · torch · tilelang · sglang-kernel · sgl-kernel · nvidia-cutlass-dsl-libs-base · nvidia-cutlass-dsl-libs-cu12 · tokenspeed-triton