$npx skillfedfor your agent

llmcompressor

A library for compressing large language models utilizing the latest techniques and research in the field for both training aware and post training techniques. The library is designed to be flexible and easy to use on top of PyTorch and HuggingFace Transformers, allowing for quick experimentation.

Worth itPyPI Software DevelopmentReleased Aug 2026221.4K downloads / moApache 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — llmcompressor-0.13.0-py3-none-any.whl
v0.13.0 · released 2026-08-11 · Python >=3.10 · 13 runtime deps: loguru, pyyaml, numpy, requests, tqdm, torch, transformers, datasets

Yes. Active maintenance, no known vulnerabilities, permissive license, and low install friction make it a solid choice for anyone deploying LLMs with vLLM. The substantial dependency footprint (torch, transformers, etc.) is expected for this use case. Start with the step-by-step compression guide in the documentation to select an appropriate quantization scheme for your model and hardware.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10+.
  • torch and transformers must be installed; llmcompressor will not function without them.
  • GPU recommended for practical compression workflows.

License · maintenance · safety

Apache 2.0 (permissive) — Apache 2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for production deployment scenarios.

last release 2026-08-11 (3 days) · last repo commit 2026-08-14 · 3,678 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 221,408 downloads/mo, #9,280 on PyPI

Verify before relying

pip install llmcompressor

from llmcompressor.transformers import SparseAutoModelForCausalLM
from transformers import AutoTokenizer

model = SparseAutoModelForCausalLM.from_pretrained(
    "RedHatAI/Llama-2-7b-chat-hf-W4A8-GPTQ"
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-chat-hf")
inputs = tokenizer("Hello world", return_tensors="pt")
outputs = model.generate(**inputs, max_length=50)
  • Whether quantized models maintain accuracy on downstream tasks beyond the calibration set used during compression.
  • Performance gains and memory savings for specific model sizes and quantization schemes in your target hardware.
  • Compatibility with vLLM versions beyond what the fact sheet documents.
Same gist for agents: .md · .json

What it is and what it does

llmcompressor is a PyTorch-based library for compressing large language models through quantization, pruning, and related techniques. It integrates with Hugging Face models and outputs compressed checkpoints in the compressed-tensors format, which vLLM can load directly for inference. The library supports multiple quantization precisions (int8, fp8, NVFP4, MXFP4, etc.) and algorithms (GPTQ, AWQ, SmoothQuant, AutoRound, REAP), with built-in support for weight-only, weight-activation, KV cache, and attention quantization.

Typical usage involves loading a Hugging Face model, applying a compression recipe (via YAML configuration or Python API), and saving the result for deployment. The library handles distributed training (DDP) and disk offloading to compress very large models on limited hardware. It depends on torch, transformers, datasets, accelerate, and several specialized packages like auto-round and compressed-tensors.

Use it for

  • Reduce model size and memory footprint for single-GPU deployment of large models like Llama or Qwen variants.
  • Apply post-training quantization (PTQ) to existing checkpoints without retraining, using calibration data.
  • Compress Mixture-of-Experts models by pruning less-relevant experts while maintaining accuracy.
  • Prepare quantized models for vLLM inference with guaranteed format compatibility.
  • Experiment with different quantization schemes (W4A8, W8A16, NVFP4, etc.) on custom models.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Active maintenance, no known vulnerabilities, permissive license, and low install friction make it a solid choice for anyone deploying LLMs with vLLM. The substantial dependency footprint (torch, transformers, etc.) is expected for this use case. Start with the step-by-step compression guide in the documentation to select an appropriate quantization scheme for your model and hardware.

Install

llmcompressor on PyPI

Before you install

Low friction install with a pure-Python wheel. Active maintenance with a recent release (3 days old) and 3678 repository stars. Requires 13 runtime dependencies including torch, transformers, and datasets—a substantial but standard ML stack.

Requires Python 3.10+. torch and transformers must be installed; llmcompressor will not function without them. GPU recommended for practical compression workflows.

License in practice

Apache 2.0 permissive license allows commercial and private use with minimal restrictions, making it suitable for production deployment scenarios.

Quickstart

pip install llmcompressor

from llmcompressor.transformers import SparseAutoModelForCausalLM
from transformers import AutoTokenizer

model = SparseAutoModelForCausalLM.from_pretrained(
    "RedHatAI/Llama-2-7b-chat-hf-W4A8-GPTQ"
)
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-7b-chat-hf")
inputs = tokenizer("Hello world", return_tensors="pt")
outputs = model.generate(**inputs, max_length=50)

Verify before relying

  • Whether quantized models maintain accuracy on downstream tasks beyond the calibration set used during compression.
  • Performance gains and memory savings for specific model sizes and quantization schemes in your target hardware.
  • Compatibility with vLLM versions beyond what the fact sheet documents.

Package facts

LicenseApache 2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
13 packages
logurupyyamlnumpyrequeststqdmtorchtransformersdatasetsauto-roundacceleratenvidia-ml-pypillowcompressed-tensors
MaintenanceActively maintained 3 days since the last release
Last repo commit
First released
Downloads221,408 / month, #9,280 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: POSIX :: LinuxProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyTopic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: MathematicsTopic :: Software DevelopmentTopic :: Software Development :: Libraries :: Python Modules

Evidence: llmcompressor-0.13.0-py3-none-any.whl

Tags

Capabilities
llm quantizationmodel compression pytorchvllm optimizationweight quantizationlanguage model pruningactivation quantizationmodel efficiency inferencehuggingface model compression
Topics
model-optimizationquantizationinference-acceleration
PyPI keywords
llmcompressorllmslarge language modelstransformerspytorchhuggingfacecompressorscompressionquantizationpruningsparsityoptimizationmodel optimizationmodel compression

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “vllm optimization”

  • llmcompressorllmcompressor optimizes large language models for efficient…
  • vllmvLLM is a high-throughput inference and serving engine for large…
  • ipex-llmAccelerates large language model inference on Intel hardware (GPU,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Software Development packages

typing-extensions Worth it
PyPI · Software Development · released Jul 2026

Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.

PSF-2.0pure Python · 3.9+
1.9Bdownloads / mo
numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
annotated-doc With conditions
PyPI · Software Development · released Jul 2026

Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.

MITpure Python · 3.9+
456.2Mdownloads / mo
typer Worth it
PyPI · Software Development · released Aug 2026

Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.

Install it if you are building CLIs in Python.

MITpure Python · 3.10+
369.3Mdownloads / mo
distlib With conditions
PyPI · Software Development · released Jun 2026

Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.

permissive licensepure Python
323.3Mdownloads / mo

See also auto-gptq · auto-round · torchao · vllm · compressed-tensors · nvidia-modelopt · vllm-tpu · mlx-lm · lm-format-enforcer · mineru-vl-utils

Further reading