$npx skillfedfor your agent

flash-attn-4

Flash Attention CUTE (CUDA Template Engine) implementation

With conditionsPyPI Artificial IntelligenceReleased Aug 20262.0M downloads / moBSD 3-Clause LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — flash_attn_4-4.0.0b26-py3-none-any.whl
v4.0.0b26 · released 2026-08-12 · Python >=3.10 · 7 runtime deps: nvidia-cutlass-dsl, torch, einops, typing_extensions, apache-tvm-ffi, torch-c-dlpack-ext, quack-kernels

Yes, if you are training or running transformer models on Hopper or Blackwell GPUs and want lower latency and memory usage. The package is actively maintained, has no known vulnerabilities, and uses a permissive license. However, it is in alpha status and only useful for specific GPU hardware; it will not benefit users on other GPU architectures or CPU-only setups.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires an NVIDIA Hopper or Blackwell GPU and corresponding CUDA toolkit; torch must be installed and configured for your CUDA version.
  • Installation is straightforward with low friction.
  • The package is actively maintained with a recent release (2 days old) and high repository engagement (24706 stars).

License · maintenance · safety

BSD 3-Clause License (permissive) — BSD 3-Clause License is permissive, allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.

last release 2026-08-12 (2 days) · last repo commit 2026-08-14 · 24,706 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,040,273 downloads/mo, #3,347 on PyPI

Verify before relying

pip install flash-attn-4
# or for CUDA 13:
# pip install "flash-attn-4[cu13]"

from flash_attn.cute import flash_attn_func
out = flash_attn_func(q, k, v, causal=True)
  • Performance benchmarks compared to standard PyTorch attention or other implementations.
  • Supported tensor shapes, dtypes, and edge cases for q, k, v inputs.
  • Memory overhead and actual speedup on different model sizes and batch configurations.
  • Stability and numerical accuracy guarantees relative to standard attention.
Same gist for agents: .md · .json

What it is and what it does

flash-attn-4 is a specialized GPU kernel library that reimplements the attention operation—a core component of transformer neural networks—using CuTeDSL for modern NVIDIA Hopper and Blackwell architectures. It trades general-purpose compatibility for speed and memory efficiency on these specific GPUs by fusing multiple attention computation steps into a single kernel, reducing memory bandwidth and improving cache locality.

The package is designed for researchers and practitioners building or fine-tuning large language models and other transformer-based systems where attention computation dominates runtime. It provides two main entry points: flash_attn_func for standard attention and flash_attn_varlen_func for variable-length sequences. Installation requires torch, nvidia-cutlass-dsl, and a compatible CUDA version; the package is in active development (alpha status) and updated frequently.

Use it for

  • Accelerate training or inference of large language models on Hopper/Blackwell GPUs by replacing standard attention with optimized kernels.
  • Reduce memory consumption during transformer model training by fusing attention operations into a single GPU kernel.
  • Implement causal attention for autoregressive language generation with lower latency.
  • Handle variable-length sequences efficiently in batched transformer inference without padding overhead.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are training or running transformer models on Hopper or Blackwell GPUs and want lower latency and memory usage.

The package is actively maintained, has no known vulnerabilities, and uses a permissive license. However, it is in alpha status and only useful for specific GPU hardware; it will not benefit users on other GPU architectures or CPU-only setups.

Install

flash-attn-4 on PyPI

Before you install

Installation is straightforward with low friction. The package is actively maintained with a recent release (2 days old) and high repository engagement (24706 stars). Requires torch and NVIDIA GPU libraries as runtime dependencies.

Requires an NVIDIA Hopper or Blackwell GPU and corresponding CUDA toolkit; torch must be installed and configured for your CUDA version.

License in practice

BSD 3-Clause License is permissive, allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.

Quickstart

pip install flash-attn-4
# or for CUDA 13:
# pip install "flash-attn-4[cu13]"

from flash_attn.cute import flash_attn_func
out = flash_attn_func(q, k, v, causal=True)

Verify before relying

  • Performance benchmarks compared to standard PyTorch attention or other implementations.
  • Supported tensor shapes, dtypes, and edge cases for q, k, v inputs.
  • Memory overhead and actual speedup on different model sizes and batch configurations.
  • Stability and numerical accuracy guarantees relative to standard attention.

Package facts

LicenseBSD 3-Clause License permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
nvidia-cutlass-dsltorcheinopstyping_extensionsapache-tvm-ffitorch-c-dlpack-extquack-kernels
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads2,040,273 / month, #3,347 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaLicense :: OSI Approved :: BSD LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12

Evidence: flash_attn_4-4.0.0b26-py3-none-any.whl

Tags

Capabilities
gpu accelerated attention mechanismhopper blackwell gpu optimizationtransformer attention kernelcuda attention optimizationefficient attention computation
Topics
gpu-kernelstransformer-optimizationcuda

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “gpu accelerated attention mechanism”

  • flash-attn-4GPU-accelerated attention mechanism implementation using CuTeDSL for…
  • flash-attnProvides optimized GPU implementations of scaled dot-product…
  • sglangSGLang is a serving framework that runs large language models and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also flash-attn · fa3-fwd · nvidia-cudnn-frontend · ring-flash-attn · flashinfer-cubin · flashinfer-python · tokenspeed-mla · transformer-engine-cu12 · transformer-engine · transformer-engine-cu13

Further reading