Gespeichert in:
| Hauptverfasser: | Zhao, Bohan, Wang, Yuanhong, Liu, Chenglin, Pan, Jiagi, Yang, Guang, Liu, Ruitao, Zhang, Tingrui, Luo, Kai, Xu, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2512.03644 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
von: Liu, Ruitao, et al.
Veröffentlicht: (2026)
Varuna: Enabling Failure-Type Aware RDMA Failover
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026)
Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
von: Bansal, Mayank, et al.
Veröffentlicht: (2026)
von: Bansal, Mayank, et al.
Veröffentlicht: (2026)
On the Resilience of Fast Failover Routing Against Dynamic Link Failures
von: Dai, Wenkai, et al.
Veröffentlicht: (2024)
von: Dai, Wenkai, et al.
Veröffentlicht: (2024)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
Seq1F1B: Efficient Sequence-Level Pipeline Parallelism for Large Language Model Training
von: Sun, Ao, et al.
Veröffentlicht: (2024)
von: Sun, Ao, et al.
Veröffentlicht: (2024)
SIMPLE: Disaggregating Sampling from GPU Inference into a Decision Plane for Faster Distributed LLM Serving
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
SiPipe: Bridging the CPU-GPU Utilization Gap for Efficient Pipeline-Parallel LLM Inference
von: He, Yongchao, et al.
Veröffentlicht: (2025)
von: He, Yongchao, et al.
Veröffentlicht: (2025)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Broadcast in Almost Mixing Time
von: Paramonov, Anton, et al.
Veröffentlicht: (2025)
von: Paramonov, Anton, et al.
Veröffentlicht: (2025)
Mell: Memory-Efficient Large Language Model Serving via Multi-GPU KV Cache Management
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
von: Qianli, Liu, et al.
Veröffentlicht: (2025)
Fast Iterative Graph Computing with Updated Neighbor States
von: Zhou, Yijie, et al.
Veröffentlicht: (2024)
von: Zhou, Yijie, et al.
Veröffentlicht: (2024)
An Almost Tight Lower Bound for Plurality Consensus with Undecided State Dynamics in the Population Protocol Model
von: El-Hayek, Antoine, et al.
Veröffentlicht: (2025)
von: El-Hayek, Antoine, et al.
Veröffentlicht: (2025)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
von: Sakip, Akhmed, et al.
Veröffentlicht: (2026)
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
von: Lin, Xinyuan, et al.
Veröffentlicht: (2025)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
von: Liu, Ziming, et al.
Veröffentlicht: (2024)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model Training
von: Han, Shujie, et al.
Veröffentlicht: (2026)
von: Han, Shujie, et al.
Veröffentlicht: (2026)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
λScale: Enabling Fast Scaling for Serverless Large Language Model Inference
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
von: Yu, Minchen, et al.
Veröffentlicht: (2025)
W4A16 Mixed-Precision Matrix Multiplication on Decoupled Architecture: Kernel Design and Memory Bottleneck Analysis for Ascend NPUs
von: He, Yuanhong, et al.
Veröffentlicht: (2026)
von: He, Yuanhong, et al.
Veröffentlicht: (2026)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
von: Zhang, Zili, et al.
Veröffentlicht: (2024)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
von: Guo, Cong, et al.
Veröffentlicht: (2024)
von: Guo, Cong, et al.
Veröffentlicht: (2024)
Sparse Checkpointing for Fast and Reliable MoE Training
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
von: Gandhi, Swapnil, et al.
Veröffentlicht: (2024)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
von: Duan, Jiangfei, et al.
Veröffentlicht: (2024)
FLeeC: a Fast Lock-Free Application Cache
von: Costa, André J., et al.
Veröffentlicht: (2024)
von: Costa, André J., et al.
Veröffentlicht: (2024)
CascadeInfer: Length-Aware Scheduling of LLM Serving with Low Latency and Load Balancing
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
von: Yuan, Yitao, et al.
Veröffentlicht: (2025)
Fast State Restoration in LLM Serving with HCache
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
von: Gao, Shiwei, et al.
Veröffentlicht: (2024)
Mangrove: Fast and Parallelizable State Replication for Blockchains
von: Paramonov, Anton, et al.
Veröffentlicht: (2025)
von: Paramonov, Anton, et al.
Veröffentlicht: (2025)
SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
von: Zhao, Yihao, et al.
Veröffentlicht: (2025)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
An Engineering Journey Training Large Language Models at Scale on Alps: The Apertus Experience
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
von: Coles, Jonathan, et al.
Veröffentlicht: (2026)
An Explorative Study on Distributed Computing Techniques in Training and Inference of Large Language Models
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
von: Hakim, Sheikh Azizul, et al.
Veröffentlicht: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
von: Dong, Xianzhe, et al.
Veröffentlicht: (2025)
FastMPS: Revisit Data Parallel in Large-scale Matrix Product State Sampling
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
von: Chen, Yaojian, et al.
Veröffentlicht: (2025)
PipeBoost: Resilient Pipelined Architecture for Fast Serverless LLM Scaling
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
von: Liu, Chongpeng, et al.
Veröffentlicht: (2025)
Building State Machine Replication Using Practical Network Synchrony
von: Wan, Yiliang, et al.
Veröffentlicht: (2025)
von: Wan, Yiliang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025) -
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
von: Liu, Ruitao, et al.
Veröffentlicht: (2026) -
Varuna: Enabling Failure-Type Aware RDMA Failover
von: Wang, Xiaoyang, et al.
Veröffentlicht: (2026) -
Uber's Failover Architecture: Reconciling Reliability and Efficiency in Hyperscale Microservice Infrastructure
von: Bansal, Mayank, et al.
Veröffentlicht: (2026) -
On the Resilience of Fast Failover Routing Against Dynamic Link Failures
von: Dai, Wenkai, et al.
Veröffentlicht: (2024)