Scaling State-Space Models on Multiple GPUs with Tensor Parallelism
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dutt, Anurag, Shah, Nimit, Masarani, Hazem, Gandhi, Anshul |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
von: Dutt, Anurag, et al.
Veröffentlicht: (2025)
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
von: Pan, Feng, et al.
Veröffentlicht: (2026)
von: Pan, Feng, et al.
Veröffentlicht: (2026)
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024)
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)
ShardTensor: Domain Parallelism for Scientific Machine Learning
von: Adams, Corey, et al.
Veröffentlicht: (2026)
von: Adams, Corey, et al.
Veröffentlicht: (2026)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
von: Hagemann, Johannes, et al.
Veröffentlicht: (2023)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
von: Singh, Siddharth, et al.
Veröffentlicht: (2023)
SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
von: Skiadopoulos, Athinagoras, et al.
Veröffentlicht: (2025)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
von: Liu, Dennis, et al.
Veröffentlicht: (2025)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
von: Chen, Ling, et al.
Veröffentlicht: (2026)
von: Chen, Ling, et al.
Veröffentlicht: (2026)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
von: Wu, Yongji, et al.
Veröffentlicht: (2025)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
von: Zhang, Xin, et al.
Veröffentlicht: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
von: Zhu, Ruidong, et al.
Veröffentlicht: (2025)
HetCCL: Accelerating LLM Training with Heterogeneous GPUs
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
von: Kim, Heehoon, et al.
Veröffentlicht: (2026)
GOGH: Correlation-Guided Orchestration of GPUs in Heterogeneous Clusters
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
von: Raeisi, Ahmad, et al.
Veröffentlicht: (2025)
LLMQ: Efficient Lower-Precision Pretraining for Consumer GPUs
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
von: Schultheis, Erik, et al.
Veröffentlicht: (2025)
Efficient Training on Multiple Consumer GPUs with RoundPipe
von: Luo, Yibin, et al.
Veröffentlicht: (2026)
von: Luo, Yibin, et al.
Veröffentlicht: (2026)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
von: Niu, Yifan, et al.
Veröffentlicht: (2026)
Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
von: Hu, Tiancheng, et al.
Veröffentlicht: (2026)
Communication-Avoiding Linear Algebraic Kernel K-Means on GPUs
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
von: Bellavita, Julian, et al.
Veröffentlicht: (2026)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
von: Akinwande, Victor, et al.
Veröffentlicht: (2024)
On Optimizing the Communication of Model Parallelism
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
von: Zhuang, Yonghao, et al.
Veröffentlicht: (2022)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
von: Chang, Li-Wen, et al.
Veröffentlicht: (2024)
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2024)
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2024)
GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism
von: Polisetty, Sandeep, et al.
Veröffentlicht: (2023)
von: Polisetty, Sandeep, et al.
Veröffentlicht: (2023)
Parallelizing Maximal Clique Enumeration on GPUs
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
von: Almasri, Mohammad, et al.
Veröffentlicht: (2022)
When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry
von: Bidollahkhani, Michael, et al.
Veröffentlicht: (2026)
von: Bidollahkhani, Michael, et al.
Veröffentlicht: (2026)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
von: Yin, Ruokai, et al.
Veröffentlicht: (2025)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
von: Zhou, Yuanchang, et al.
Veröffentlicht: (2024)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
Heterogeneous Parallelism for Multimodal Large Language Model Training
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
von: Karnati, Yashaswi, et al.
Veröffentlicht: (2026)
vTensor: Flexible Virtual Tensor Management for Efficient LLM Serving
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
von: Xu, Jiale, et al.
Veröffentlicht: (2024)
Efficient Parallel Reinforcement Learning Framework using the Reactor Model
von: Kwok, Jacky, et al.
Veröffentlicht: (2023)
von: Kwok, Jacky, et al.
Veröffentlicht: (2023)
ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
von: Tahir, Anique, et al.
Veröffentlicht: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
von: Zhang, Geng, et al.
Veröffentlicht: (2025)
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2023)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2023)
Rashomon Sets and Model Multiplicity in Federated Learning
von: Heilmann, Xenia, et al.
Veröffentlicht: (2026)
von: Heilmann, Xenia, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
von: Dutt, Anurag, et al.
Veröffentlicht: (2025) -
Parallelizing Large-Scale Tensor Network Contraction on Multiple GPUs
von: Pan, Feng, et al.
Veröffentlicht: (2026) -
Nonuniform-Tensor-Parallelism: Mitigating GPU failure impact for Scaled-up LLM Training
von: Arfeen, Daiyaan, et al.
Veröffentlicht: (2025) -
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
von: Jiang, Ziheng, et al.
Veröffentlicht: (2024) -
AMPED: Accelerating MTTKRP for Billion-Scale Sparse Tensor Decomposition on Multiple GPUs
von: Wijeratne, Sasindu, et al.
Veröffentlicht: (2025)