$npx skillfedfor your agent

megatron-fsdp

**Megatron-FSDP** is an NVIDIA-developed PyTorch extension that provides a high-performance implementation of Fully Sharded Data Parallelism (FSDP)

With conditionsPyPI LibrariesReleased Jul 202676.4K downloads / moApache 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — megatron_fsdp-0.5.1-py3-none-any.whl
v0.5.1 · released 2026-07-21 · Python >=3.10 · 3 runtime deps: torch, einops, packaging

Yes, if you are training large PyTorch models on multi-GPU NVIDIA clusters and need fine-grained control over distributed sharding strategies. The package is actively maintained, has no known vulnerabilities, installs with low friction, and integrates cleanly with PyTorch's distributed ecosystem. Not necessary for single-GPU training or if you are already satisfied with standard PyTorch FSDP or other parallelism frameworks.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires torch.distributed setup and multiple GPUs; designed for data-center-scale training on NVIDIA GPUs.
  • Low friction: pure Python wheel with only torch, einops, and packaging as runtime dependencies.
  • Active maintenance with recent releases; repo shows 17428 stars and current development.

License · maintenance · safety

Apache 2.0 (permissive) — Apache 2.0 is permissive; you can use this in commercial projects, modify it, and distribute it with minimal restrictions, though you must include a copy of the license and note any changes.

last release 2026-07-21 (24 days) · last repo commit 2026-08-14 · 17,428 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 76,431 downloads/mo, #14,626 on PyPI

Verify before relying

pip install megatron-fsdp

import torch
from megatron_fsdp import fully_shard_model, fully_shard_optimizer

torch.distributed.init_process_group()
model = torch.nn.Transformer()
fsdp_model = fully_shard_model(module=model, fsdp_unit_modules=[torch.nn.TransformerEncoder])
optimizer = fully_shard_optimizer(torch.optim.AdamW(fsdp_model.parameters()))
  • Whether the package supports AMD or other non-NVIDIA GPU architectures despite NVIDIA authorship.
  • Performance benchmarks or memory savings compared to standard PyTorch FSDP or other parallelism strategies.
  • Compatibility with specific transformer frameworks beyond those listed (Megatron-Core, TransformerEngine, NeMo).
Same gist for agents: .md · .json

What it is and what it does

Megatron-FSDP is NVIDIA's PyTorch implementation of Fully Sharded Data Parallelism, a distributed training technique that splits model parameters, gradients, and optimizer states across multiple GPUs to reduce per-device memory usage. It lets you train models too large to fit on a single GPU by coordinating computation and communication across a cluster of NVIDIA GPUs. The library exposes a high-level API (fully_shard_model, fully_shard_optimizer) that wraps your model and optimizer, and supports multiple sharding strategies—from no sharding (similar to standard data parallelism) to aggressive parameter sharding (ZeRO-3 style)—so you can tune the trade-off between memory efficiency and communication overhead.

The package integrates with PyTorch's distributed checkpoint system, DeviceMesh, and DTensor abstractions, and is designed to work alongside Megatron-Core, TransformerEngine, and NVIDIA's NeMo framework. It handles the complexity of initializing extremely large models on meta-device to avoid out-of-memory errors during setup, and supports advanced parallelism patterns like tensor parallelism, context parallelism, and expert parallelism for mixture-of-experts models.

Use it for

  • Train large transformer models (billions of parameters) that cannot fit in a single GPU's memory.
  • Reduce per-device memory footprint by distributing optimizer state and gradients across a multi-GPU cluster.
  • Combine FSDP with tensor parallelism or expert parallelism for hybrid distributed training strategies.
  • Save and restore distributed checkpoints of fully-sharded models and optimizers across training runs.
  • Tune memory-communication trade-offs by switching between ZeRO-1, ZeRO-2, and ZeRO-3 sharding strategies.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are training large PyTorch models on multi-GPU NVIDIA clusters and need fine-grained control over distributed sharding strategies.

The package is actively maintained, has no known vulnerabilities, installs with low friction, and integrates cleanly with PyTorch's distributed ecosystem. Not necessary for single-GPU training or if you are already satisfied with standard PyTorch FSDP or other parallelism frameworks.

Install

megatron-fsdp on PyPI

Before you install

Low friction: pure Python wheel with only torch, einops, and packaging as runtime dependencies. Active maintenance with recent releases; repo shows 17428 stars and current development.

Requires torch.distributed setup and multiple GPUs; designed for data-center-scale training on NVIDIA GPUs.

License in practice

Apache 2.0 is permissive; you can use this in commercial projects, modify it, and distribute it with minimal restrictions, though you must include a copy of the license and note any changes.

Quickstart

pip install megatron-fsdp

import torch
from megatron_fsdp import fully_shard_model, fully_shard_optimizer

torch.distributed.init_process_group()
model = torch.nn.Transformer()
fsdp_model = fully_shard_model(module=model, fsdp_unit_modules=[torch.nn.TransformerEncoder])
optimizer = fully_shard_optimizer(torch.optim.AdamW(fsdp_model.parameters()))

Verify before relying

  • Whether the package supports AMD or other non-NVIDIA GPU architectures despite NVIDIA authorship.
  • Performance benchmarks or memory savings compared to standard PyTorch FSDP or other parallelism strategies.
  • Compatibility with specific transformer frameworks beyond those listed (Megatron-Core, TransformerEngine, NeMo).

Package facts

LicenseApache 2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
torcheinopspackaging
MaintenanceActively maintained 24 days since the last release
Last repo commit
First released
Downloads76,431 / month, #14,626 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: Image RecognitionTopic :: Scientific/Engineering :: MathematicsTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python ModulesTopic :: Utilities

Evidence: megatron_fsdp-0.5.1-py3-none-any.whl

Tags

Capabilities
distributed model training pytorchfully sharded data parallelismlarge language model parallelismmulti-gpu model shardingfsdp pytorch implementationzero redundancy optimizergpu memory efficient training
Topics
distributed-trainingmodel-parallelismgpu-optimization
PyPI keywords
NLPNLUdeepgpulanguagelearningmachinenvidiapytorchtorchtransformer

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “fully sharded data parallelism”

  • megatron-fsdpMegatron-FSDP implements Fully Sharded Data Parallelism (FSDP) in…
  • tbbProvides Python bindings to Intel's oneAPI Threading Building Blocks…
  • tbb-develProvides C++ parallelism primitives and task scheduling for…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also fairscale · megatron-core · spmd-types · torch · torchft-nightly · torchtitan · s3torchconnector · s3torchconnectorclient · aistore · flashoptim

Further reading