megatron-core
Megatron Core - a library for efficient and scalable training of transformer based models
Decision gist · record as of 2026-08-14
Yes, if you are building a custom distributed training framework or scaling transformer training across multiple GPUs. The library is production-stable, actively maintained, permissively licensed, and has no known vulnerabilities. Install friction is moderate due to GPU/CUDA requirements and torch dependency. Not recommended for simple single-GPU training or inference-only use cases.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires NVIDIA GPU(s) and torch compiled with CUDA support; building from source can consume significant memory—set MAX_JOBS environment variable if build fails.
- Medium friction: precompiled wheels available for Python 3.11 and 3.12 on x86_64 and aarch64, but requires torch and numpy.
- Package now requires Python 3.12 or later.
License · maintenance · safety
Apache 2.0 (permissive) — Apache 2.0 permissive license allows commercial use, modification, and redistribution with minimal restrictions, making it suitable for both research and production deployments.
last release 2026-07-21 (24 days) · last repo commit 2026-08-14 · 17,427 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 208,041 downloads/mo, #9,536 on PyPI
Alternatives
Verify before relying
pip install megatron-core
import torch
from megatron.core.models.gpt import GPTModel
from megatron.core.tensor_parallel import ColumnParallelLinear- Whether the package supports inference-only workflows or is primarily designed for training pipelines.
- Compatibility and integration requirements with specific NVIDIA GPU architectures beyond H100.
- Whether checkpoint conversion with Megatron Bridge is included or requires separate installation.
What it is and what it does
Megatron Core is a composable library of GPU-optimized building blocks for training transformer models across distributed systems. It provides modular components for tensor parallelism, pipeline parallelism, data parallelism, expert parallelism, and context parallelism, along with support for mixed precision training (FP16, BF16, FP8, FP4) and model architectures including GPT, BERT, and Mamba-based models. The library is designed for framework developers and ML engineers building custom training pipelines, not as a standalone training script.
The package depends on torch, numpy, and packaging, and is actively maintained by NVIDIA with recent releases including dynamic context parallelism and multi-data center training support. Precompiled wheels are available for modern Python versions. Performance benchmarks show up to 47% Model FLOP Utilization on H100 clusters when training models from 2B to 462B parameters.
Use it for
- Build custom distributed training frameworks that need composable parallelism strategies and transformer building blocks.
- Scale transformer model training across thousands of GPUs with optimized communication and computation overlap.
- Implement mixed-precision training (FP16, BF16, FP8, FP4) for large language models to reduce memory and compute costs.
- Train variable-length sequence models using dynamic context parallelism for adaptive efficiency gains.
- Integrate advanced model architectures (GPT, BERT, Mamba) into production training pipelines with fault tolerance.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are building a custom distributed training framework or scaling transformer training across multiple GPUs.
The library is production-stable, actively maintained, permissively licensed, and has no known vulnerabilities. Install friction is moderate due to GPU/CUDA requirements and torch dependency. Not recommended for simple single-GPU training or inference-only use cases.
Install
megatron-core on PyPI
Before you install
Medium friction: precompiled wheels available for Python 3.11 and 3.12 on x86_64 and aarch64, but requires torch and numpy. Package now requires Python 3.12 or later.
Requires NVIDIA GPU(s) and torch compiled with CUDA support; building from source can consume significant memory—set MAX_JOBS environment variable if build fails.
License in practice
Apache 2.0 permissive license allows commercial use, modification, and redistribution with minimal restrictions, making it suitable for both research and production deployments.
Quickstart
pip install megatron-core
import torch
from megatron.core.models.gpt import GPTModel
from megatron.core.tensor_parallel import ColumnParallelLinear
Verify before relying
- Whether the package supports inference-only workflows or is primarily designed for training pipelines.
- Compatibility and integration requirements with specific NVIDIA GPU architectures beyond H100.
- Whether checkpoint conversion with Megatron Bridge is included or requires separate installation.
Package facts
| License | Apache 2.0 permissive |
| Python support | Supports the current Python release >=3.12 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 3 packagestorchnumpypackaging |
| Maintenance | Actively maintained 24 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 208,041 / month, #9,536 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Image RecognitionTopic :: Scientific/Engineering :: MathematicsTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python ModulesTopic :: Utilities |
Evidence: megatron_core-0.18.2-cp311-cp311-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; megatron_core-0.18.2-cp311-cp311-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; megatron_core-0.18.2-cp312-cp312-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; megatron_core-0.18.2-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; megatron_core-0.18.2-cp313-cp313-manylinux_2_24_aarch64.manylinux_2_28_aarch64.whl; megatron_core-0.18.2-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “distributed transformer training”
- megatron-coreMegatron Core provides GPU-optimized building blocks and parallelism…
- transformer-engine-cu12Accelerates Transformer model training and inference on NVIDIA GPUs…
- transformer-engine-cu13Accelerates Transformer model training and inference on NVIDIA GPUs…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also megatron-fsdp · transformer-engine-cu12 · transformer-engine · transformer-engine-cu13 · fairscale · deepspeed · xformers · torchtitan · spmd-types · nvidia-modelopt