$npx skillfedfor your agent

sgl-kernel

Kernel Library for SGLang

With conditionsPyPI Artificial IntelligenceReleased Jan 2026162.7K downloads / mopermissive licensePlatform wheel

Decision gist · record as of 2026-08-14

platform wheels — sgl_kernel-0.3.21-cp310-abi3-manylinux2014_aarch64.whl · sgl_kernel-0.3.21-cp310-abi3-manylinux2014_x86_64.whl
v0.3.21 · released 2026-01-14 · Python >=3.10

Yes, if you are building or extending an LLM inference engine and your environment uses torch == 2.9.1. The library is actively maintained, has no Python runtime dependencies, and is already proven in production LLM systems. The strict torch version requirement is a real constraint—verify compatibility before committing. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires torch == 2.9.1 and Python ≥3.10; NVIDIA GPU with CUDA support required for runtime execution.
  • Medium install friction due to compiled wheel distribution (x86_64 and aarch64 manylinux2014 wheels available).
  • Requires torch == 2.9.1 as a hard dependency.

License · maintenance · safety

permissive license (permissive) — Apache License 2.0 is permissive; you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license and attribution notices.

last release 2026-01-14 (212 days) · last repo commit 2026-08-14 · 31,815 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 162,742 downloads/mo, #10,590 on PyPI

Verify before relying

pip install sgl-kernel

import sgl_kernel
# Use optimized kernels via sgl_kernel module in your inference engine
  • Whether torch == 2.9.1 is pinned strictly or if compatible newer/older versions work in practice
  • GPU compute capability requirements beyond what the description implies
  • Performance characteristics relative to other LLM inference kernel libraries
Same gist for agents: .md · .json

What it is and what it does

sgl-kernel is a compiled CUDA kernel library designed to accelerate large language model inference by providing optimized compute primitives. It exposes custom kernel operations that LLM inference engines like SGLang and LightLLM use to speed up matrix operations and other compute-intensive tasks during model execution. The package ships as pre-built wheels for x86_64 and aarch64 architectures, requiring Python ≥3.10 and torch == 2.9.1.

The library is intended for developers building or extending LLM inference systems who need lower-level kernel control and optimization. It has no Python runtime dependencies beyond torch, making it a thin wrapper around CUDA code. Installation is straightforward via pip, though the strict torch version requirement means you must align your environment carefully. The project is actively maintained with recent commits and has accumulated significant adoption in the LLM inference ecosystem.

Use it for

  • Accelerate matrix operations in custom LLM inference engines by replacing generic CUDA kernels with optimized primitives.
  • Build vision-language model inference pipelines that require specialized compute kernels for efficient token processing.
  • Integrate into SGLang or LightLLM deployments to improve throughput and latency of model serving.
  • Develop custom LLM serving frameworks that need fine-grained control over GPU compute without writing raw CUDA code.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are building or extending an LLM inference engine and your environment uses torch == 2.9.1.

The library is actively maintained, has no Python runtime dependencies, and is already proven in production LLM systems. The strict torch version requirement is a real constraint—verify compatibility before committing. No known security vulnerabilities.

Install

sgl-kernel on PyPI

Before you install

Medium install friction due to compiled wheel distribution (x86_64 and aarch64 manylinux2014 wheels available). Requires torch == 2.9.1 as a hard dependency. Active maintenance with recent commits and no runtime Python dependencies simplifies integration once the torch version constraint is satisfied.

Requires torch == 2.9.1 and Python ≥3.10; NVIDIA GPU with CUDA support required for runtime execution.

License in practice

Apache License 2.0 is permissive; you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license and attribution notices.

Quickstart

pip install sgl-kernel

import sgl_kernel
# Use optimized kernels via sgl_kernel module in your inference engine

Verify before relying

  • Whether torch == 2.9.1 is pinned strictly or if compatible newer/older versions work in practice
  • GPU compute capability requirements beyond what the description implies
  • Performance characteristics relative to other LLM inference kernel libraries

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.10
Install frictionMedium. Platform-specific wheel
Runtime dependenciesNone
MaintenanceActively maintained 212 days since the last release
Last repo commit
First released
Downloads162,742 / month, #10,590 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Environment :: GPU :: NVIDIA CUDALicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3

Evidence: sgl_kernel-0.3.21-cp310-abi3-manylinux2014_aarch64.whl; sgl_kernel-0.3.21-cp310-abi3-manylinux2014_x86_64.whl

Tags

Capabilities
llm inference kernelscuda optimization for language modelscustom gpu kernels llmefficient inference primitivesvision-language model accelerationsglang kernel libraryllm compute optimization
Topics
cuda-kernelsllm-inferencegpu-acceleration

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “custom gpu kernels llm”

  • sgl-kernelsgl-kernel provides optimized CUDA compute primitives for LLM…
  • sglang-kernelsglang-kernel provides optimized CUDA compute primitives for LLM…
  • flashinfer-pythonFlashInfer provides optimized GPU kernels for LLM inference…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also liger-kernel · sgl-deep-gemm · sglang-kernel · sglang · mineru-vl-utils · flashinfer-python · apache-tvm-ffi · fa3-fwd · kernels-data · smg-grpc-servicer

Further reading