gram-newton-schulz
Fast Newton-Schulz Algorithm with Kernels
Decision gist · record as of 2026-08-14
Yes, if you are training on NVIDIA Hopper or Blackwell GPUs with PyTorch 2.7.1+ and use or plan to use Muon as your optimizer. The package is actively maintained, has no known vulnerabilities, carries permissive MIT licensing, and offers a mathematically equivalent but faster polar decomposition with no training accuracy tradeoff. No, if you lack the required GPU hardware or are locked to older PyTorch versions—the installation is hardware-specific and will fail gracefully without the target GPU.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA Hopper (H100) or Blackwell (B200/B300) GPU, PyTorch 2.7.1+, and CUDA 12.9+.
- Must install PyTorch first before running pip install with --no-build-isolation.
- Low friction install via pip with a pure-Python wheel, but requires NVIDIA Hopper or Blackwell GPU hardware, PyTorch 2.7.1+, and CUDA 12.9+.
License · maintenance · safety
MIT (permissive) — MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
last release 2026-07-02 (43 days) · last repo commit 2026-07-02 · 174 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 176,350 downloads/mo, #10,246 on PyPI
Alternatives
Verify before relying
pip install gram-newton-schulz --no-build-isolation
from gram_newton_schulz import GramNewtonSchulz, POLAR_EXPRESS_COEFFICIENTS
gram_NS = GramNewtonSchulz(
ns_coefficients=POLAR_EXPRESS_COEFFICIENTS,
gram_newton_schulz_reset_iterations=[2]
)
result = gram_NS(X)- Actual speedup magnitude relative to standard Newton-Schulz in real training scenarios beyond the claimed 2x in the description.
- Numerical stability guarantees and convergence behavior across different matrix shapes and conditioning.
- Compatibility with torch.compile on Hopper GPUs and workaround effectiveness on Blackwell.
What it is and what it does
Gram Newton-Schulz is a specialized optimizer component that accelerates polar decomposition—a key operation in the Muon optimizer—by iterating on a smaller symmetric Gram matrix instead of the full rectangular matrix. This mathematical reformulation reduces floating-point operations and enables more efficient symmetric GEMM kernels on modern NVIDIA GPUs. The package provides both a standalone GramNewtonSchulz callable class and integration into a full Muon optimizer that handles mixed parameter types (2D weights for orthogonalization, scalars via an auxiliary optimizer) with autotuned restart points for numerical stability.
It depends on torch for tensor operations, quack-kernels for custom GEMM implementations, and nvidia-cutlass-dsl for low-level GPU kernel generation. The algorithm is mathematically equivalent to standard Newton-Schulz with no claimed training accuracy loss, making it a direct swap for existing Muon-based training pipelines. Installation requires explicit GPU hardware (Hopper or Blackwell) and careful dependency management to avoid pip installing a CPU-only PyTorch.
Use it for
- Drop-in replacement for Newton-Schulz in Muon optimizer for faster large-model training on H100 or B200/B300 clusters.
- Accelerating polar decomposition in custom orthogonalization-based optimizers that need symmetric matrix operations.
- Autotuning restart schedules for Newton-Schulz iterations to balance speed and numerical stability in iterative algorithms.
- Training transformer models with Muon where 2D weight matrices benefit from symmetric GEMM kernels on Hopper/Blackwell hardware.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are training on NVIDIA Hopper or Blackwell GPUs with PyTorch 2.7.1+ and use or plan to use Muon as your optimizer.
The package is actively maintained, has no known vulnerabilities, carries permissive MIT licensing, and offers a mathematically equivalent but faster polar decomposition with no training accuracy tradeoff. No, if you lack the required GPU hardware or are locked to older PyTorch versions—the installation is hardware-specific and will fail gracefully without the target GPU.
Install
gram-newton-schulz on PyPI
Before you install
Low friction install via pip with a pure-Python wheel, but requires NVIDIA Hopper or Blackwell GPU hardware, PyTorch 2.7.1+, and CUDA 12.9+. Must use --no-build-isolation flag to avoid pip installing a CPU-only PyTorch. Active maintenance with recent commits.
Requires NVIDIA Hopper (H100) or Blackwell (B200/B300) GPU, PyTorch 2.7.1+, and CUDA 12.9+. Must install PyTorch first before running pip install with --no-build-isolation.
License in practice
MIT license is permissive and imposes no restrictions on use, modification, or distribution in proprietary or open-source projects.
Quickstart
pip install gram-newton-schulz --no-build-isolation
from gram_newton_schulz import GramNewtonSchulz, POLAR_EXPRESS_COEFFICIENTS
gram_NS = GramNewtonSchulz(
ns_coefficients=POLAR_EXPRESS_COEFFICIENTS,
gram_newton_schulz_reset_iterations=[2]
)
result = gram_NS(X)
Verify before relying
- Actual speedup magnitude relative to standard Newton-Schulz in real training scenarios beyond the claimed 2x in the description.
- Numerical stability guarantees and convergence behavior across different matrix shapes and conditioning.
- Compatibility with torch.compile on Hopper GPUs and workaround effectiveness on Blackwell.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagestorchquack-kernelsnvidia-cutlass-dsl |
| Maintenance | Actively maintained 43 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 176,350 / month, #10,246 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: gram_newton_schulz-0.1.6-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “polar decomposition pytorch”
- gram-newton-schulzImplements a hardware-aware Gram Newton-Schulz algorithm for polar…
- pytorch-waveletsProvides 2D discrete wavelet and dual-tree complex wavelet transforms…
- tensorlyTensorLy performs tensor decomposition, tensor learning, and tensor…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also sgl-deep-gemm · nvidia-cudnn-frontend · nvidia-cublas · nvidia-cusolver-cu11 · transformer-engine-cu12 · nvidia-cusparse-cu12 · nvidia-cutlass-dsl-libs-base · nvidia-cutlass-dsl-libs-core · pyroots · phonors