flash-attn
Flash Attention: Fast and Memory-Efficient Exact Attention
Decision gist · record as of 2026-08-14
Yes, if you have a supported GPU (Ampere, Ada, Hopper for CUDA; MI200/MI300 for ROCm) and are training or deploying transformer models where attention is a bottleneck. The high install friction (compilation, GPU toolkit requirements) is justified by significant speed and memory gains. Not worth installing for CPU-only workflows or unsupported GPU architectures (e.g., older Turing GPUs without FlashAttention 1.x).AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA 12.0+ or ROCm 6.0+, PyTorch 2.2+, ninja build tool, and Linux (Windows support experimental as of v2.3.2).
- Compilation requires a GPU and appropriate toolkit installed.
- High install friction: requires CUDA 12.0+ or ROCm 6.0+, PyTorch 2.2+, ninja build tool, and compilation from source.
License · maintenance · safety
permissive license (permissive) — BSD License (permissive): you may use, modify, and distribute the package freely, including in commercial projects, provided you retain the license notice and credit the authors as requested in the repository.
last release 2026-06-11 (64 days) · last repo commit 2026-08-14 · 24,706 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 484,138 downloads/mo, #6,407 on PyPI
Alternatives
Verify before relying
pip install flash-attn --no-build-isolation
from flash_attn import flash_attn_func
# flash_attn_func(q, k, v, dropout_p=0.0, softmax_scale=None, causal=False)- Exact performance gains over standard attention on specific hardware and sequence lengths
- Compatibility with all transformer architectures and attention variants beyond those documented
- Windows support stability and completeness as of current version
What it is and what it does
FlashAttention is a GPU-accelerated library that implements fast, memory-efficient exact attention mechanisms for transformer models. It provides optimized kernels for scaled dot-product attention that reduce memory I/O and improve computational efficiency on modern GPUs. The package includes FlashAttention-2 (optimized for Ampere, Ada, and Hopper architectures) and a beta FlashAttention-3 for Hopper GPUs, with support for both NVIDIA CUDA and AMD ROCm backends.
The library integrates with PyTorch and einops, exposing functions like flash_attn_func and flash_attn_qkvpacked_func that drop in as replacements for standard attention. It supports various attention patterns including causal masking, sliding window attention, multi-query and grouped query attention, dropout, rotary embeddings, and ALiBi. Installation requires compilation against your GPU toolkit, which adds setup complexity but is necessary to generate hardware-specific optimized code.
Use it for
- Accelerate training and inference of large language models and transformers on NVIDIA A100, RTX 4090, H100, or AMD MI200/MI300 GPUs
- Reduce memory consumption during attention computation, enabling longer sequences or larger batch sizes on the same hardware
- Implement causal attention for autoregressive language model generation with improved efficiency
- Deploy transformer models with lower latency by replacing standard PyTorch attention with optimized FlashAttention kernels
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you have a supported GPU (Ampere, Ada, Hopper for CUDA; MI200/MI300 for ROCm) and are training or deploying transformer models where attention is a bottleneck.
The high install friction (compilation, GPU toolkit requirements) is justified by significant speed and memory gains. Not worth installing for CPU-only workflows or unsupported GPU architectures (e.g., older Turing GPUs without FlashAttention 1.x).
Install
flash-attn on PyPI
Before you install
High install friction: requires CUDA 12.0+ or ROCm 6.0+, PyTorch 2.2+, ninja build tool, and compilation from source. Compilation takes 3–5 minutes on well-equipped machines but can exhaust RAM on systems with many CPU cores and <96GB memory; MAX_JOBS environment variable can mitigate this. Package is actively maintained with recent releases.
Requires CUDA 12.0+ or ROCm 6.0+, PyTorch 2.2+, ninja build tool, and Linux (Windows support experimental as of v2.3.2). Compilation requires a GPU and appropriate toolkit installed.
License in practice
BSD License (permissive): you may use, modify, and distribute the package freely, including in commercial projects, provided you retain the license notice and credit the authors as requested in the repository.
Quickstart
pip install flash-attn --no-build-isolation
from flash_attn import flash_attn_func
# flash_attn_func(q, k, v, dropout_p=0.0, softmax_scale=None, causal=False)
Verify before relying
- Exact performance gains over standard attention on specific hardware and sequence lengths
- Compatibility with all transformer architectures and attention variants beyond those documented
- Windows support stability and completeness as of current version
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | High. Source build required |
| Runtime dependencies | 2 packagestorcheinops |
| Maintenance | Actively maintained 64 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 484,138 / month, #6,407 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: BSD LicenseOperating System :: UnixProgramming Language :: Python :: 3 |
Evidence: flash_attn-2.8.3.post1.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “fast attention mechanism GPU”
- flash-attnProvides optimized GPU implementations of scaled dot-product…
- flash-attn-4GPU-accelerated attention mechanism implementation using CuTeDSL for…
- CoLT5-attentionImplements conditionally routed efficient attention mechanisms for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also flash-attn-4 · ring-flash-attn · fa3-fwd · tokenspeed-mla · sageattention · local-attention · flashinfer-cubin · flashinfer-python · mamba-ssm · CoLT5-attention