cut-cross-entropy
Code for cut cross entropy, a memory efficient implementation of linear-cross-entropy loss.
Decision gist · record as of 2026-08-14
Yes, with conditions. Install if you train large-vocabulary language models on memory-constrained GPUs and need dramatic memory reduction in the loss computation layer. The package is well-motivated by published research and offers both optimized Triton kernels and fallback implementations. However, verify the license terms before production use (license treatment is unclear), and be aware that maintenance is dormant—no active development is expected. Requires Python 3.10+, PyTorch 2.4+, and Triton 3.0+ on Ampere+ GPUs.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10+, PyTorch 2.4+, Triton 3.0+, and Ampere or newer GPU.
- Triton is not available on macOS; the package falls back to torch.compile on unsupported platforms.
- Low friction install with pure Python wheel distribution.
License · maintenance · safety
(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available in the metadata. Verify the actual license terms before using in production or proprietary work.
last release 2025-01-07 (584 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 609,745 downloads/mo, #5,770 on PyPI
Alternatives
Verify before relying
pip install cut-cross-entropy
from cut_cross_entropy import linear_cross_entropy
embeddings = model.compute_embedding(inputs)
classifier = model.get_classifier_weights()
loss = linear_cross_entropy(embeddings, classifier, labels)- Actual license terms and restrictions (license_treatment is unclear in metadata)
- Whether dormant maintenance status indicates ongoing support or abandonment
- Compatibility with PyTorch and Triton versions beyond those explicitly mentioned
What it is and what it does
Cut Cross-Entropy (CCE) is a memory-efficient implementation of the cross-entropy loss computation for training large-vocabulary language models. During standard training, the cross-entropy layer materializes a full logit matrix covering all token-vocabulary pairs, consuming enormous amounts of GPU memory—often more than the rest of the model combined. CCE avoids this by computing logits on-the-fly in flash memory, computing only the logit for the correct token and evaluating log-sum-exp over all vocabulary items without materializing the full matrix.
The package provides a drop-in replacement function `linear_cross_entropy` that accepts embeddings, classifier weights, and labels, with optional support for token-shifting in causal language modeling. It includes custom Triton kernels for Ampere and newer GPUs, a torch.compile fallback for older hardware and non-Linux platforms, and direct integration patches for transformers models (Llama, Phi3, Mistral, Gemma2 families). According to the documentation, this reduces memory consumption from 24 GB to 1 MB for the loss computation on Gemma 2 (2B), with negligible impact on training speed.
Use it for
- Training large language models on GPUs with limited memory by reducing classifier head memory footprint from tens of gigabytes to under a gigabyte.
- Fine-tuning transformer models (Llama, Phi3, Mistral, Gemma2) using the transformers library without modifying model code via cce_patch.
- Computing per-token loss and perplexity efficiently with reduction='none' for detailed loss analysis without materializing full logit matrices.
- Enabling training on older GPUs or non-Linux systems via torch.compile fallback when Triton kernels are unavailable.
- Reducing overall training-time memory consumption of the classifier head to fit larger models or batch sizes on fixed hardware.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
Install if you train large-vocabulary language models on memory-constrained GPUs and need dramatic memory reduction in the loss computation layer. The package is well-motivated by published research and offers both optimized Triton kernels and fallback implementations. However, verify the license terms before production use (license treatment is unclear), and be aware that maintenance is dormant—no active development is expected. Requires Python 3.10+, PyTorch 2.4+, and Triton 3.0+ on Ampere+ GPUs.
Install
cut-cross-entropy on PyPI
Before you install
Low friction install with pure Python wheel distribution. Maintenance status is dormant (584 days since release), though the package received a recent update on 2025-01-07. Depends on torch and triton, both widely available.
Requires Python 3.10+, PyTorch 2.4+, Triton 3.0+, and Ampere or newer GPU. Triton is not available on macOS; the package falls back to torch.compile on unsupported platforms.
License in practice
License treatment is unclear—no SPDX identifier or raw license text is available in the metadata. Verify the actual license terms before using in production or proprietary work.
Quickstart
pip install cut-cross-entropy
from cut_cross_entropy import linear_cross_entropy
embeddings = model.compute_embedding(inputs)
classifier = model.get_classifier_weights()
loss = linear_cross_entropy(embeddings, classifier, labels)
Verify before relying
- Actual license terms and restrictions (license_treatment is unclear in metadata)
- Whether dormant maintenance status indicates ongoing support or abandonment
- Compatibility with PyTorch and Triton versions beyond those explicitly mentioned
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagestorchtriton |
| Maintenance | Dormant 584 days since the last release |
| First released | |
| Downloads | 609,745 / month, #5,770 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: cut_cross_entropy-25.1.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “memory efficient cross entropy loss”
- cut-cross-entropyComputes cross-entropy loss for large-vocabulary language models with…
- tomesdSpeeds up Stable Diffusion image generation by merging redundant…
- liger-kernelLiger Kernel provides optimized Triton kernels for LLM training,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also entmax · pytorch-metric-learning · pyctcdecode · liger-kernel · sentence-transformers · sgl-kernel · torch · transformer-smaller-training-vocab · rax · coqui-tts-trainer