Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Arfeen, Daiyaan, Mudigere, Dheevatsa, More, Ankit, Gopireddy, Bhargava, Inci, Ahmet, Ganger, Gregory R. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2024)
by: Arfeen, Daiyaan, et al.
Published: (2024)
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025)
by: Bhatia, Nidhi, et al.
Published: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
by: Chen, Shawn Shuoshuo, et al.
Published: (2025)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
by: Liang, Antian, et al.
Published: (2025)
by: Liang, Antian, et al.
Published: (2025)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
by: Li, Shengwei, et al.
Published: (2023)
by: Li, Shengwei, et al.
Published: (2023)
cuFastTuckerPlus: A Stochastic Parallel Sparse FastTucker Decomposition Using GPU Tensor Cores
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Synergistic Tensor and Pipeline Parallelism
by: Qi, Mengshi, et al.
Published: (2025)
by: Qi, Mengshi, et al.
Published: (2025)
Large Scale Multi-GPU Based Parallel Traffic Simulation for Accelerated Traffic Assignment and Propagation
by: Jiang, Xuan, et al.
Published: (2024)
by: Jiang, Xuan, et al.
Published: (2024)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
by: Ai, Xin, et al.
Published: (2024)
by: Ai, Xin, et al.
Published: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2025)
by: Wijeratne, Sasindu, et al.
Published: (2025)
Accelerating Drug Discovery in AutoDock-GPU with Tensor Cores
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
by: Tang, Ding, et al.
Published: (2024)
by: Tang, Ding, et al.
Published: (2024)
Speeding up Local Optimization in Vehicle Routing with Tensor-based GPU Acceleration
by: Lei, Zhenyu, et al.
Published: (2025)
by: Lei, Zhenyu, et al.
Published: (2025)
GPU-Based Parallel Computing Methods for Medical Photoacoustic Image Reconstruction
by: Yi, Xinyao, et al.
Published: (2024)
by: Yi, Xinyao, et al.
Published: (2024)
Concurrent Scheduling of High-Level Parallel Programs on Multi-GPU Systems
by: Knorr, Fabian, et al.
Published: (2025)
by: Knorr, Fabian, et al.
Published: (2025)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
by: Mitra, Tiyasa, et al.
Published: (2025)
by: Mitra, Tiyasa, et al.
Published: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
by: Liu, Man, et al.
Published: (2026)
by: Liu, Man, et al.
Published: (2026)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
by: Chen, Haoyu, et al.
Published: (2025)
by: Chen, Haoyu, et al.
Published: (2025)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
by: Wang, Zhigang, et al.
Published: (2024)
by: Wang, Zhigang, et al.
Published: (2024)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
by: Srivatsa, Vikranth, et al.
Published: (2026)
by: Srivatsa, Vikranth, et al.
Published: (2026)
Characterization of GPU TEE Overheads in Distributed Data Parallel ML Training
by: Lee, Jonghyun, et al.
Published: (2025)
by: Lee, Jonghyun, et al.
Published: (2025)
Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
by: Mo, Zizhao, et al.
Published: (2025)
by: Mo, Zizhao, et al.
Published: (2025)
Multi-GPU Acceleration of PALABOS Fluid Solver using C++ Standard Parallelism
by: Latt, Jonas, et al.
Published: (2025)
by: Latt, Jonas, et al.
Published: (2025)
Heimdall++: Optimizing GPU Utilization and Pipeline Parallelism for Efficient Single-Pulse Detection
by: Xia, Bingzheng, et al.
Published: (2025)
by: Xia, Bingzheng, et al.
Published: (2025)
Parallel GPU-Enabled Algorithms for SpGEMM on Arbitrary Semirings with Hybrid Communication
by: McFarland, Thomas, et al.
Published: (2025)
by: McFarland, Thomas, et al.
Published: (2025)
Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
by: Dutt, Anurag, et al.
Published: (2026)
by: Dutt, Anurag, et al.
Published: (2026)
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
by: Pan, Feng, et al.
Published: (2026)
by: Pan, Feng, et al.
Published: (2026)
Minimizing Communication for Parallel Symmetric Tensor Times Same Vector Computation
by: Daas, Hussam Al, et al.
Published: (2025)
by: Daas, Hussam Al, et al.
Published: (2025)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
by: Xia, Mengchun, et al.
Published: (2026)
by: Xia, Mengchun, et al.
Published: (2026)
APEX: Asynchronous Parallel CPU-GPU Execution for Online LLM Inference on Constrained GPUs
by: Fan, Jiakun, et al.
Published: (2025)
by: Fan, Jiakun, et al.
Published: (2025)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
A Unified CPU-GPU Protocol for GNN Training
by: Lin, Yi-Chien, et al.
Published: (2024)
by: Lin, Yi-Chien, et al.
Published: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
by: Shyam, Vasu, et al.
Published: (2026)
by: Shyam, Vasu, et al.
Published: (2026)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
by: Xu, Wendong, et al.
Published: (2025)
by: Xu, Wendong, et al.
Published: (2025)
Neutron particle transport 3D method of characteristic Multi GPU platform Parallel Computing
by: Zhou, Faguo, et al.
Published: (2025)
by: Zhou, Faguo, et al.
Published: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
by: He, Yongchao, et al.
Published: (2025)
by: He, Yongchao, et al.
Published: (2025)
Similar Items
-
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
by: Arfeen, Daiyaan, et al.
Published: (2024) -
Helix Parallelism: Rethinking Sharding Strategies for Interactive Multi-Million-Token LLM Decoding
by: Bhatia, Nidhi, et al.
Published: (2025) -
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024) -
Toward Co-adapting Machine Learning Job Shape and Cluster Topology
by: Chen, Shawn Shuoshuo, et al.
Published: (2025) -
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
by: Liang, Antian, et al.
Published: (2025)