humming-kernels
Quantization GEMM Kernel
Decision gist · record as of 2026-08-14
Yes, if you need quantized GEMM kernels for inference on NVIDIA GPUs and have the required hardware (SM75+) and Python environment (≥3.10). The package is actively maintained, has no known vulnerabilities, and offers broad quantization format support. However, verify the license terms before use, and confirm that your CUDA and PyTorch versions are compatible—the fact sheet does not specify exact version constraints beyond the architecture requirement.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU with SM75+ architecture (Turing or newer); NVCC compiler and CUDA toolkit must be available; Python >=3.10.
- Medium install friction due to 9 runtime dependencies including torch, triton, and cuda-bindings.
- Package is actively maintained with recent releases, but availability is limited to manylinux wheels for x86_64 and aarch64 architectures.
License · maintenance · safety
(unclear) — License status is unclear—no SPDX identifier or raw license text is available. Verify the project's actual license terms before use in proprietary or redistributed code.
last release 2026-07-31 (14 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,421,523 downloads/mo, #3,923 on PyPI
Alternatives
Verify before relying
pip install humming-kernels
import torch
from humming.layer import HummingLayer
layer = HummingLayer(
shape_n=8192,
shape_k=8192,
weight_config={"dtype": "int6"},
torch_dtype=torch.float16,
).cuda()
weight = torch.randn((8192, 8192), dtype=torch.float16, device="cuda:0")
inputs = torch.randn((128, 8192), dtype=torch.float16, device="cuda:0")
layer.load_from_unquantized(weight)
layer.transform()
output = layer(inputs)- Whether the package's license permits commercial use and redistribution—license_treatment is unclear.
- Performance benchmarks and throughput claims relative to other quantized GEMM libraries.
- Exact CUDA and PyTorch version compatibility constraints beyond the stated SM architecture requirements.
What it is and what it does
Humming is a lightweight, JIT-compiled GEMM kernel library optimized for quantized inference on NVIDIA GPUs. It provides high-performance matrix multiplication kernels that support a wide range of quantization formats—from FP16 and BF16 down to FP4 and INT4 weights—paired with various activation types (FP16, BF16, FP8, INT8, INT4). The library handles both dense matrix operations and mixture-of-experts (MoE) patterns, making it suitable for deploying quantized large language models and other inference workloads.
The package is designed to be minimal and self-contained, requiring only PyTorch and NVCC as core dependencies, with a compact footprint under 100KB. It abstracts away kernel tuning through a HummingLayer interface that automatically selects appropriate kernels for your hardware and quantization configuration. You load unquantized weights, transform them into Humming's internal format, and then run inference through the layer—the library handles the low-level kernel dispatch and optimization.
Use it for
- Deploy quantized LLMs with mixed-bit weight formats (INT4, INT6, INT8) on NVIDIA GPUs for reduced memory and latency.
- Run inference with FP8 or FP4 activations and weights on newer GPUs (SM89+) for extreme compression.
- Accelerate MoE model inference by using Humming's specialized MoE GEMM kernels instead of generic matrix operations.
- Benchmark quantization strategies across different bit-widths and scale types without writing custom CUDA code.
- Integrate quantized inference into production systems where minimal dependencies and small package size are constraints.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need quantized GEMM kernels for inference on NVIDIA GPUs and have the required hardware (SM75+) and Python environment (≥3.10).
The package is actively maintained, has no known vulnerabilities, and offers broad quantization format support. However, verify the license terms before use, and confirm that your CUDA and PyTorch versions are compatible—the fact sheet does not specify exact version constraints beyond the architecture requirement.
Install
humming-kernels on PyPI
Before you install
Medium install friction due to 9 runtime dependencies including torch, triton, and cuda-bindings. Package is actively maintained with recent releases, but availability is limited to manylinux wheels for x86_64 and aarch64 architectures.
Requires NVIDIA GPU with SM75+ architecture (Turing or newer); NVCC compiler and CUDA toolkit must be available; Python >=3.10.
License in practice
License status is unclear—no SPDX identifier or raw license text is available. Verify the project's actual license terms before use in proprietary or redistributed code.
Quickstart
pip install humming-kernels
import torch
from humming.layer import HummingLayer
layer = HummingLayer(
shape_n=8192,
shape_k=8192,
weight_config={"dtype": "int6"},
torch_dtype=torch.float16,
).cuda()
weight = torch.randn((8192, 8192), dtype=torch.float16, device="cuda:0")
inputs = torch.randn((128, 8192), dtype=torch.float16, device="cuda:0")
layer.load_from_unquantized(weight)
layer.transform()
output = layer(inputs)
Verify before relying
- Whether the package's license permits commercial use and redistribution—license_treatment is unclear.
- Performance benchmarks and throughput claims relative to other quantized GEMM libraries.
- Exact CUDA and PyTorch version compatibility constraints beyond the stated SM architecture requirements.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 9 packagestorchtritonnumpysafetensorsjinja2nvidia-ml-pycuda-bindingstqdmtabulate |
| Maintenance | Actively maintained 14 days since the last release |
| First released | |
| Downloads | 1,421,523 / month, #3,923 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: humming_kernels-0.1.12-py3-none-manylinux_2_28_aarch64.whl; humming_kernels-0.1.12-py3-none-manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “quantized gemm kernels”
- humming-kernelsHumming is a JIT-compiled GEMM kernel library for quantized matrix…
- flashinfer-pythonFlashInfer provides optimized GPU kernels for LLM inference…
- nvidia-cudnn-frontendProvides Python and C++ APIs to NVIDIA's cuDNN library, exposing…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also comfy-kitchen · nvidia-cusparselt-cu12 · tokenspeed-mla · sageattention · nvidia-cudnn-frontend · sgl-deep-gemm · auto-gptq · nvidia-cusparselt-cu13 · tilelang · flashinfer-python