DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
-
Updated
Sep 19, 2026 - Python
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Making large AI models cheaper, faster and more accessible
A GPipe implementation in PyTorch
飞桨大模型开发套件,提供大语言模型、跨模态大模型、生物计算大模型等领域的全流程开发工具链。
LiBai(李白): A Toolbox for Large-Scale Distributed Parallel Training
Slicing a PyTorch Tensor Into Parallel Shards
A curated list of awesome projects and papers for distributed training or inference
Easy Parallel Library (EPL) is a general and efficient deep learning framework for distributed model training.
Distributed training (multi-node) of a Transformer model
Large scale 4D parallelism pre-training for 🤗 transformers in Mixture of Experts *(still work in progress)*
NAACL '24 (Best Demo Paper RunnerUp) / MlSys @ NeurIPS '23 - RedCoast: A Lightweight Tool to Automate Distributed Training and Inference
SC23 Deep Learning at Scale Tutorial Material
Distributed training of DNNs • C++/MPI Proxies (GPT-2, GPT-3, CosmoFlow, DLRM)
Distributed peer-to-peer LLM inference. Your prompt never leaves your device in clear text.
🧠 One LLM, split across a Mac (Apple MPS) + a Windows PC (NVIDIA CUDA) — heterogeneous pipeline inference over a framework-neutral wire, bit-for-bit identical to single-machine. No torch.distributed, no datacenter.
Halo is an open-source framework built by White Circle for training large language and multimodal models
Deep Learning at Scale Training Event at NERSC
Deep learning for science school material 2025
WIP. Veloce is a low-code Ray-based parallelization library that makes machine learning computation novel, efficient, and heterogeneous.
Deep Learning at Scale @ SC25
To associate your repository with the model-parallelism topic, visit your repo's landing page and select "manage topics."