ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
Fuente:
arXiv
Salvato in:
| Autori principali: | Won, William, Heo, Taekyung, Rashidi, Saeed, Sridharan, Srinivas, Srinivasan, Sudarshan, Krishna, Tushar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
di: Won, William, et al.
Pubblicazione: (2021)
di: Won, William, et al.
Pubblicazione: (2021)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2022)
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2022)
Towards a Standardized Representation for Deep Learning Collective Algorithms
di: Yoo, Jinsun, et al.
Pubblicazione: (2024)
di: Yoo, Jinsun, et al.
Pubblicazione: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
di: Won, William, et al.
Pubblicazione: (2023)
di: Won, William, et al.
Pubblicazione: (2023)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
di: Yoo, Jinsun, et al.
Pubblicazione: (2026)
di: Yoo, Jinsun, et al.
Pubblicazione: (2026)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
di: Raju, Aditi, et al.
Pubblicazione: (2025)
di: Raju, Aditi, et al.
Pubblicazione: (2025)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
di: Wu, Hanjiang, et al.
Pubblicazione: (2026)
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
di: Man, Changhai, et al.
Pubblicazione: (2025)
di: Man, Changhai, et al.
Pubblicazione: (2025)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
di: Cho, Jaehong, et al.
Pubblicazione: (2026)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
di: Zhang, Zili, et al.
Pubblicazione: (2024)
di: Zhang, Zili, et al.
Pubblicazione: (2024)
Lotus: Optimizing Disaggregated Transactions with Disaggregated Locks
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
di: Hu, Zhisheng, et al.
Pubblicazione: (2025)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
di: Garg, Raveesh, et al.
Pubblicazione: (2025)
di: Garg, Raveesh, et al.
Pubblicazione: (2025)
RollMux: Phase-Level Multiplexing for Disaggregated RL Post-Training
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
di: Wu, Tianyuan, et al.
Pubblicazione: (2025)
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
di: Lai, Ruiqi, et al.
Pubblicazione: (2025)
ScalePool: Hybrid XLink-CXL Fabric for Composable Resource Disaggregation in Unified Scale-up Domains
di: Woo, Hyein, et al.
Pubblicazione: (2025)
di: Woo, Hyein, et al.
Pubblicazione: (2025)
INDIGO: Page Migration for Hardware Memory Disaggregation Across a Network
di: Patke, Archit, et al.
Pubblicazione: (2025)
di: Patke, Archit, et al.
Pubblicazione: (2025)
HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
di: Dong, Xianzhe, et al.
Pubblicazione: (2025)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
di: Sridharan, Srinivas, et al.
Pubblicazione: (2026)
TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-Scale
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
di: Yoon, Dongha, et al.
Pubblicazione: (2025)
DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
di: Basit, Omar, et al.
Pubblicazione: (2026)
di: Basit, Omar, et al.
Pubblicazione: (2026)
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
di: Zhong, Yinmin, et al.
Pubblicazione: (2024)
InfiniLoRA: Disaggregated Multi-LoRA Serving for Large Language Models
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
di: Chen, Hongyu, et al.
Pubblicazione: (2026)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
di: Niam, Arefin, et al.
Pubblicazione: (2025)
di: Niam, Arefin, et al.
Pubblicazione: (2025)
Disaggregated Prefill and Decoding Inference System for Large Language Model Serving on Multi-Vendor GPUs
di: Chen, Xing, et al.
Pubblicazione: (2025)
di: Chen, Xing, et al.
Pubblicazione: (2025)
Software Resource Disaggregation for HPC with Serverless Computing
di: Copik, Marcin, et al.
Pubblicazione: (2024)
di: Copik, Marcin, et al.
Pubblicazione: (2024)
LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation
di: Li, Mufei, et al.
Pubblicazione: (2024)
di: Li, Mufei, et al.
Pubblicazione: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
di: Puri, Amit, et al.
Pubblicazione: (2023)
di: Puri, Amit, et al.
Pubblicazione: (2023)
From Attention to Disaggregation: Tracing the Evolution of LLM Inference
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
di: Kumar, Madabattula Rajesh, et al.
Pubblicazione: (2025)
Towards Disaggregation-Native Data Streaming between Devices
di: Asmussen, Nils, et al.
Pubblicazione: (2024)
di: Asmussen, Nils, et al.
Pubblicazione: (2024)
SWARM: Replicating Shared Disaggregated-Memory Data in No Time
di: Murat, Antoine, et al.
Pubblicazione: (2024)
di: Murat, Antoine, et al.
Pubblicazione: (2024)
Optimizing Distributed Training Approaches for Scaling Neural Networks
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
di: Zhang, Hanze, et al.
Pubblicazione: (2025)
Arrow: Adaptive Scheduling Mechanisms for Disaggregated LLM Inference Architecture
di: Wu, Yu, et al.
Pubblicazione: (2025)
di: Wu, Yu, et al.
Pubblicazione: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
di: He, Wenhao, et al.
Pubblicazione: (2026)
di: He, Wenhao, et al.
Pubblicazione: (2026)
Proceedings of 3rd Workshop on Heterogeneous Composable and Disaggregated Systems
di: Pinto, Christian, et al.
Pubblicazione: (2024)
di: Pinto, Christian, et al.
Pubblicazione: (2024)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
di: Zhang, Zhexiang, et al.
Pubblicazione: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
di: Yarlagadda, Srihas, et al.
Pubblicazione: (2025)
di: Yarlagadda, Srihas, et al.
Pubblicazione: (2025)
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
di: Zhu, Ruidong, et al.
Pubblicazione: (2025)
Taming Request Imbalance: SLO-Aware Scheduling for Disaggregated LLM Inference
di: Wang, Qipeng
Pubblicazione: (2026)
di: Wang, Qipeng
Pubblicazione: (2026)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
di: Ruan, Chaoyi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
di: Won, William, et al.
Pubblicazione: (2021) -
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2022) -
Towards a Standardized Representation for Deep Learning Collective Algorithms
di: Yoo, Jinsun, et al.
Pubblicazione: (2024) -
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
di: Won, William, et al.
Pubblicazione: (2023) -
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
di: Yoo, Jinsun, et al.
Pubblicazione: (2026)