Echo: Simulating Distributed Training At Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Feng, Yicheng, Chen, Yuetao, Chen, Kaiwen, Li, Jingzong, Wu, Tianyuan, Cheng, Peng, Wu, Chuan, Wang, Wei, Ho, Tsung-Yi, Xu, Hong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
von: Xue, Chunyu, et al.
Veröffentlicht: (2026)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
von: Golden, Alicia, et al.
Veröffentlicht: (2025)
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)
von: Sun, Ao, et al.
Veröffentlicht: (2025)
DSV: Exploiting Dynamic Sparsity to Accelerate Large-Scale Video DiT Training
von: Tan, Xin, et al.
Veröffentlicht: (2025)
von: Tan, Xin, et al.
Veröffentlicht: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2024)
On-the-fly Communication-and-Computing to Enable Representation Learning for Distributed Point Clouds
von: Chen, Xu, et al.
Veröffentlicht: (2024)
von: Chen, Xu, et al.
Veröffentlicht: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
QSync: Quantization-Minimized Synchronous Distributed Training Across Hybrid Devices
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
von: Zhao, Juntao, et al.
Veröffentlicht: (2024)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
von: Zhang, Shiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Shiwei, et al.
Veröffentlicht: (2024)
MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
von: Zhao, Juntao, et al.
Veröffentlicht: (2025)
Accelerating Distributed MoE Training and Inference with Lina
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
von: Li, Jiamin, et al.
Veröffentlicht: (2022)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
von: Chen, Liangkun, et al.
Veröffentlicht: (2025)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2023)
DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
von: Tian, Ye, et al.
Veröffentlicht: (2024)
von: Tian, Ye, et al.
Veröffentlicht: (2024)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Wu, Tianyuan, et al.
Veröffentlicht: (2025)
Optimizing Distributed Training Approaches for Scaling Neural Networks
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
von: Baligodugula, Vishnu Vardhan, et al.
Veröffentlicht: (2025)
DeepServe: Serverless Large Language Model Serving at Scale
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
von: Hu, Junhao, et al.
Veröffentlicht: (2025)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
von: Wang, Yanbo, et al.
Veröffentlicht: (2026)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinxiao, et al.
Veröffentlicht: (2026)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
A Study on the Performance of Distributed Training of Data-driven CFD Simulations
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
von: Iserte, Sergio, et al.
Veröffentlicht: (2026)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
von: Sun, Mo, et al.
Veröffentlicht: (2024)
von: Sun, Mo, et al.
Veröffentlicht: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
von: Zhuang, Chen, et al.
Veröffentlicht: (2024)
CaraServe: CPU-Assisted and Rank-Aware LoRA Serving for Generative LLM Inference
von: Li, Suyi, et al.
Veröffentlicht: (2024)
von: Li, Suyi, et al.
Veröffentlicht: (2024)
Communication-Efficient Sparsely-Activated Model Training via Sequence Migration and Token Condensation
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
von: Chen, Fahao, et al.
Veröffentlicht: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
von: Shen, Tao, et al.
Veröffentlicht: (2025)
von: Shen, Tao, et al.
Veröffentlicht: (2025)
ModTrans: Translating Real-world Models for Distributed Training Simulator
von: Lyu, Yi
Veröffentlicht: (2026)
von: Lyu, Yi
Veröffentlicht: (2026)
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
von: Qiu, Zhimin, et al.
Veröffentlicht: (2025)
von: Qiu, Zhimin, et al.
Veröffentlicht: (2025)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
von: Lu, Kuan-Wei, et al.
Veröffentlicht: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
von: Svedas, Jonas, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024) -
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
von: Xue, Chunyu, et al.
Veröffentlicht: (2026) -
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
von: Yang, Yi, et al.
Veröffentlicht: (2025) -
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
von: Golden, Alicia, et al.
Veröffentlicht: (2025) -
BurstEngine: an Efficient Distributed Framework for Training Transformers on Extremely Long Sequences of over 1M Tokens
von: Sun, Ao, et al.
Veröffentlicht: (2025)