$npx skillfedfor your agent

deepspeed

DeepSpeed library

With conditionsPyPI Artificial IntelligenceReleased Aug 20261.3M downloads / moApache Software License 2.0Source build

Decision gist · record as of 2026-08-14

sdist only — deepspeed-0.19.5.tar.gz · builds from source
v0.19.5 · released 2026-08-10 · 11 runtime deps: einops, hjson, msgpack, ninja, numpy, packaging, psutil, py-cpuinfo

Yes, if you are training large models on multi-GPU or multi-node clusters. The high install friction and compiled dependencies are justified by the substantial memory and speed gains for distributed training at scale. Not necessary for single-GPU training of small models. Active maintenance, permissive license, and zero known vulnerabilities support adoption.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires torch and ninja to be installed; compilation of native extensions during setup; GPU or multi-node hardware strongly recommended for practical use.
  • High install friction: requires torch, ninja, and several compiled dependencies (einops, msgpack, psutil, py-cpuinfo).
  • Active maintenance with recent release (4 days old) and substantial community adoption (42930 GitHub stars, 1.2M monthly downloads).

License · maintenance · safety

Apache Software License 2.0 (permissive) — Apache Software License 2.0 is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most production and research contexts.

last release 2026-08-10 (4 days) · last repo commit 2026-08-14 · 42,930 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,265,457 downloads/mo, #4,143 on PyPI

Verify before relying

pip install deepspeed torch ninja
import deepspeed
model_engine, optimizer, _, _ = deepspeed.initialize(model=model, model_parameters=model.parameters(), config_params=ds_config)
  • Specific performance gains or scaling limits for different model sizes and hardware configurations.
  • Compatibility matrix with specific torch versions beyond Python version classifiers.
  • Whether all 11 runtime dependencies are mandatory or some are optional for certain features.
Same gist for agents: .md · .json

What it is and what it does

DeepSpeed is a distributed training framework built on top of PyTorch that enables efficient training of very large language models by combining multiple system-level optimizations. It implements techniques like ZeRO (Zero Redundancy Optimizer) for memory efficiency, gradient checkpointing, and offloading to CPU or NVMe storage, allowing models that would otherwise exceed GPU memory to train on available hardware. The library integrates with popular frameworks including Transformers, Accelerate, Lightning, and others, and has been used to train models ranging from billions to hundreds of billions of parameters.

The package is actively maintained by Microsoft's AI at Scale initiative and has been central to training some of the largest open-source language models. It requires torch as a core dependency along with build tools (ninja) and system utilities (psutil, py-cpuinfo). Installation has high friction due to compiled components, but the library is designed for multi-GPU and multi-node distributed training scenarios where that overhead is negligible compared to training time.

Use it for

  • Train large language models (billions of parameters) that exceed single GPU memory by distributing computation and offloading intermediate states.
  • Reduce training time for existing models through gradient accumulation, mixed precision, and efficient communication patterns across multiple GPUs.
  • Fine-tune pretrained models on limited hardware by leveraging memory optimization techniques like ZeRO without rewriting training loops.
  • Implement custom distributed training pipelines with automatic parallelism strategies via configuration rather than code changes.
  • Integrate distributed training into existing PyTorch workflows via Transformers or Accelerate without major refactoring.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are training large models on multi-GPU or multi-node clusters.

The high install friction and compiled dependencies are justified by the substantial memory and speed gains for distributed training at scale. Not necessary for single-GPU training of small models. Active maintenance, permissive license, and zero known vulnerabilities support adoption.

Install

deepspeed on PyPI

Before you install

High install friction: requires torch, ninja, and several compiled dependencies (einops, msgpack, psutil, py-cpuinfo). Active maintenance with recent release (4 days old) and substantial community adoption (42930 GitHub stars, 1.2M monthly downloads).

Requires torch and ninja to be installed; compilation of native extensions during setup; GPU or multi-node hardware strongly recommended for practical use.

License in practice

Apache Software License 2.0 is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most production and research contexts.

Quickstart

pip install deepspeed torch ninja
import deepspeed
model_engine, optimizer, _, _ = deepspeed.initialize(model=model, model_parameters=model.parameters(), config_params=ds_config)

Verify before relying

  • Specific performance gains or scaling limits for different model sizes and hardware configurations.
  • Compatibility matrix with specific torch versions beyond Python version classifiers.
  • Whether all 11 runtime dependencies are mandatory or some are optional for certain features.

Package facts

LicenseApache Software License 2.0 permissive
Python supportNot specified
Install frictionHigh. Source build required
Runtime dependencies
11 packages
einopshjsonmsgpackninjanumpypackagingpsutilpy-cpuinfopydantictorchtqdm
MaintenanceActively maintained 4 days since the last release
Last repo commit
First released
Downloads1,265,457 / month, #4,143 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: deepspeed-0.19.5.tar.gz

Tags

Capabilities
distributed deep learning traininglarge language model training optimizationGPU memory efficient trainingzero redundancy optimizermodel parallelism framework
Topics
distributed-traininglarge-language-modelsgpu-optimization

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “GPU memory efficient training”

  • deepspeedDeepSpeed is a distributed deep learning training library that…
  • unsloth-zooUnsloth Zoo provides utilities for fine-tuning large language models…
  • unslothUnsloth accelerates training and fine-tuning of large language models…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also transformer-engine · torchtitan · megatron-core · accelerate · nvidia-cudnn-cu11 · transformer-engine-cu12 · sglang · memcache-hybrid · skypilot · skypilot-nightly

Further reading