causal-conv1d
Causal depthwise conv1d in CUDA, with a PyTorch interface
Decision gist · record as of 2026-08-14
Yes, if you are building or fine-tuning models that use causal convolutions and need lower latency than PyTorch's standard conv1d. The high install friction (CUDA, build tools, ROCm patching for some users) is a real cost; install only if you have a GPU environment already set up and the performance gain justifies the build complexity. No known vulnerabilities and active maintenance are positive signals.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires CUDA toolkit, C++ compiler, and ninja build tool.
- ROCm 6.0 users must apply the rocm6_0.patch before installation.
- Requires Python ≥3.9.
License · maintenance · safety
permissive license (permissive) — BSD License (permissive) places no restrictions on use, modification, or redistribution in proprietary or open-source projects.
last release 2026-05-09 (97 days) · last repo commit 2026-08-14 · 936 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 121,015 downloads/mo, #12,000 on PyPI
Alternatives
Verify before relying
pip install causal-conv1d
from causal_conv1d import causal_conv1d_fn
# x: (batch, dim, seqlen), weight: (dim, width), bias: (dim,)
out = causal_conv1d_fn(x, weight, bias=None, activation=None)- Whether the package supports AMD GPUs via ROCm beyond the documented 6.0 and 6.1 versions.
- Performance benchmarks or latency comparisons against torch.nn.functional.conv1d with equivalent padding.
- Whether activation parameter accepts custom functions or only the documented 'silu'/'swish' strings.
What it is and what it does
Causal-conv1d provides a CUDA kernel for depthwise 1D convolution that maintains causality—the output at each position depends only on current and past inputs, never future ones. It wraps this kernel with a PyTorch interface, accepting batched sequences and returning convolved outputs in the same shape. The operation is mathematically equivalent to torch.nn.functional.conv1d with causal padding, but implemented as a fused CUDA kernel for lower latency and memory overhead.
The package targets machine learning workloads where causal convolutions are essential, such as autoregressive language models, time-series processing, and streaming inference. It supports mixed precision (fp32, fp16, bf16) and optional SiLU/Swish activation. Installation requires a working CUDA environment and build tools; ROCm users on version 6.0 must patch their installation first.
Use it for
- Accelerating causal convolution layers in transformer-based language models during training and inference.
- Implementing efficient streaming time-series models that process sequences incrementally without lookahead.
- Reducing latency in real-time audio or signal processing pipelines that depend on causal filtering.
- Mixed-precision inference on edge GPUs where fp16 or bf16 precision reduces memory and bandwidth.
- Building state-space models (e.g., Mamba, S4) that use causal convolutions as a core operation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building or fine-tuning models that use causal convolutions and need lower latency than PyTorch's standard conv1d.
The high install friction (CUDA, build tools, ROCm patching for some users) is a real cost; install only if you have a GPU environment already set up and the performance gain justifies the build complexity. No known vulnerabilities and active maintenance are positive signals.
Install
causal-conv1d on PyPI
Before you install
High install friction: requires torch, packaging, and ninja as runtime dependencies, plus a C++ compiler and CUDA toolchain. ROCm users on version 6.0 need to apply a patch to avoid compilation errors. Active maintenance with recent commits, but the build-from-source requirement makes installation non-trivial.
Requires CUDA toolkit, C++ compiler, and ninja build tool. ROCm 6.0 users must apply the rocm6_0.patch before installation. Requires Python ≥3.9.
License in practice
BSD License (permissive) places no restrictions on use, modification, or redistribution in proprietary or open-source projects.
Quickstart
pip install causal-conv1d
from causal_conv1d import causal_conv1d_fn
# x: (batch, dim, seqlen), weight: (dim, width), bias: (dim,)
out = causal_conv1d_fn(x, weight, bias=None, activation=None)
Verify before relying
- Whether the package supports AMD GPUs via ROCm beyond the documented 6.0 and 6.1 versions.
- Performance benchmarks or latency comparisons against torch.nn.functional.conv1d with equivalent padding.
- Whether activation parameter accepts custom functions or only the documented 'silu'/'swish' strings.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | High. Source build required |
| Runtime dependencies | 3 packagestorchpackagingninja |
| Maintenance | Actively maintained 97 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 121,015 / month, #12,000 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: BSD LicenseOperating System :: UnixProgramming Language :: Python :: 3 |
Evidence: causal_conv1d-1.6.2.post1.tar.gz
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “causal convolution pytorch cuda”
- causal-conv1dImplements a CUDA-optimized causal depthwise 1D convolution operation…
- juliusJulius provides differentiable, GPU-accelerated digital signal…
- flash-attnProvides optimized GPU implementations of scaled dot-product…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also cuequivariance-ops-torch-cu12 · nvidia-cudnn-frontend · nvidia-cusparselt-cu12 · comfy-kitchen · resize-right · nvidia-cusparselt-cu13 · unfoldNd · julius · fa3-fwd · transformer-engine-cu12