flydsl
FlyDSL - ROCm Domain Specific Language for layout algebra (Python + embedded MLIR runtime)
Decision gist · record as of 2026-08-14
Yes, if you are developing GPU kernels for ROCm hardware and want to express layouts and tiling at a high level. The package is actively maintained, has no known vulnerabilities, and bundles MLIR so setup is simpler than standalone MLIR. Install friction is medium due to ROCm dependency and optional source builds, but pre-built wheels for Python 3.10–3.14 are available. Not suitable if you need CUDA, OpenCL, or non-ROCm GPU support.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- ROCm is required for GPU execution, tests, and benchmarks; Python 3.8+ required; source builds need cmake >=3.20 and C++17 compiler.
- Medium install friction: wheels are available for Python 3.10–3.14 on Linux x86_64, but the package bundles MLIR bindings and requires ROCm for GPU execution and testing.
- Source builds demand LLVM/MLIR compilation (30+ minutes) and C++17 tooling.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and proprietary use with minimal restrictions; attribution required.
last release 2026-08-09 (5 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 107,165 downloads/mo, #12,627 on PyPI
Alternatives
Verify before relying
pip install flydsl
import flydsl
print('FlyDSL installed')- Whether the embedded MLIR runtime is fully functional without a separate MLIR Python wheel installation in all environments.
- Performance characteristics and optimization potential compared to other GPU kernel DSLs or hand-written kernels.
- Maturity and stability of the layout algebra system and compiler passes in production use.
What it is and what it does
FlyDSL is a Python domain-specific language for writing GPU kernels that run on ROCm hardware. It sits atop an embedded MLIR compiler stack (the Fly dialect) and lets you express kernel structure, data layouts, tiling strategies, and memory movement at a high level in Python, then compile them down to GPU machine code. The package bundles MLIR Python bindings so you don't need a separate MLIR installation.
The core abstraction is a layout system—Shape, Stride, and Layout objects that map logical coordinates to physical memory indices. You compose these layouts using algebra operations (composition, product, partition) to express complex data access patterns like swizzling, tiling, and vectorization. You decorate Python functions with `@flyc.kernel` or `@flyc.jit`, write kernel logic using FlyDSL's expression API (arithmetic, vector, GPU, ROCDL, buffer, math, and memory operations), and the JIT compiler handles lowering to GPU code. Pre-built kernels for GEMM, MoE, Softmax, and Norm are included.
Use it for
- Author custom GPU kernels with explicit tiling and memory layout control without writing CUDA or HIP directly.
- Prototype and tune high-performance matrix multiplication and tensor operations on ROCm GPUs.
- Express complex data movement and swizzling patterns using layout algebra instead of manual indexing.
- Build production GPU kernels (GEMM, MoE, normalization) with the included pre-built kernel library.
- Benchmark and profile kernel performance using the built-in autotune and benchmarking infrastructure.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are developing GPU kernels for ROCm hardware and want to express layouts and tiling at a high level.
The package is actively maintained, has no known vulnerabilities, and bundles MLIR so setup is simpler than standalone MLIR. Install friction is medium due to ROCm dependency and optional source builds, but pre-built wheels for Python 3.10–3.14 are available. Not suitable if you need CUDA, OpenCL, or non-ROCm GPU support.
Install
flydsl on PyPI
Before you install
Medium install friction: wheels are available for Python 3.10–3.14 on Linux x86_64, but the package bundles MLIR bindings and requires ROCm for GPU execution and testing. Source builds demand LLVM/MLIR compilation (30+ minutes) and C++17 tooling.
ROCm is required for GPU execution, tests, and benchmarks; Python 3.8+ required; source builds need cmake >=3.20 and C++17 compiler.
License in practice
Apache-2.0 permissive license allows commercial and proprietary use with minimal restrictions; attribution required.
Quickstart
pip install flydsl
import flydsl
print('FlyDSL installed')
Verify before relying
- Whether the embedded MLIR runtime is fully functional without a separate MLIR Python wheel installation in all environments.
- Performance characteristics and optimization potential compared to other GPU kernel DSLs or hand-written kernels.
- Maturity and stability of the layout algebra system and compiler passes in production use.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 5 days since the last release |
| First released | |
| Downloads | 107,165 / month, #12,627 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: flydsl-0.3.1-cp310-cp310-manylinux_2_27_x86_64.whl; flydsl-0.3.1-cp311-cp311-manylinux_2_27_x86_64.whl; flydsl-0.3.1-cp312-cp312-manylinux_2_27_x86_64.whl; flydsl-0.3.1-cp313-cp313-manylinux_2_27_x86_64.whl; flydsl-0.3.1-cp314-cp314-manylinux_2_27_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “gpu kernel dsl python”
- flydslFlyDSL is a Python DSL and embedded MLIR compiler for authoring…
- tilelangTilelang is a domain-specific language for writing high-performance…
- nvidia-cutlass-dsl-libs-cu13Provides a Python DSL for writing high-performance CUDA kernels using…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also nvidia-cutlass-dsl-libs-base · xdsl · nvidia-cutlass-dsl-libs-cu13 · nvidia-cutlass-dsl · nvidia-cutlass-dsl-libs-cu12 · nvidia-cutlass-dsl-libs-core · tilelang · pyrsmi · sglang-kernel · tosa-tools