TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Won, William, Elavazhagan, Midhilesh, Srinivasan, Sudarshan, Gupta, Swati, Krishna, Tushar |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
by: Bambhaniya, Abhimanyu, et al.
Published: (2024)
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026)
by: Wu, Hanjiang, et al.
Published: (2026)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)
Towards a Standardized Representation for Deep Learning Collective Algorithms
by: Yoo, Jinsun, et al.
Published: (2024)
by: Yoo, Jinsun, et al.
Published: (2024)
Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
Heterogeneity-Aware Client Selection Methodology For Efficient Federated Learning
by: Balivada, Nihal, et al.
Published: (2026)
by: Balivada, Nihal, et al.
Published: (2026)
When Less is More: Achieving Faster Convergence in Distributed Edge Machine Learning
by: Basani, Advik Raj, et al.
Published: (2024)
by: Basani, Advik Raj, et al.
Published: (2024)
Algorithms for Collaborative Machine Learning under Statistical Heterogeneity
by: Hahn, Seok-Ju
Published: (2024)
by: Hahn, Seok-Ju
Published: (2024)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
LACS: Learning-Augmented Algorithms for Carbon-Aware Resource Scaling with Uncertain Demand
by: Bostandoost, Roozbeh, et al.
Published: (2024)
by: Bostandoost, Roozbeh, et al.
Published: (2024)
CAFE: Carbon-Aware Federated Learning in Geographically Distributed Data Centers
by: Bian, Jieming, et al.
Published: (2023)
by: Bian, Jieming, et al.
Published: (2023)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Efficient Distributed Learning over Decentralized Networks with Convoluted Support Vector Machine
by: Chen, Canyi, et al.
Published: (2025)
by: Chen, Canyi, et al.
Published: (2025)
D3FL: Data Distribution and Detrending for Robust Federated Learning in Non-linear Time-series Data
by: Marisetty, Harsha Varun, et al.
Published: (2025)
by: Marisetty, Harsha Varun, et al.
Published: (2025)
StraightLine: An End-to-End Resource-Aware Scheduler for Machine Learning Application Requests
by: Ching, Cheng-Wei, et al.
Published: (2024)
by: Ching, Cheng-Wei, et al.
Published: (2024)
DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems
by: Dai, Yuanjun, et al.
Published: (2025)
by: Dai, Yuanjun, et al.
Published: (2025)
Priority-Aware Model-Distributed Inference at Edge Networks
by: Li, Teng, et al.
Published: (2024)
by: Li, Teng, et al.
Published: (2024)
FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing Maps
by: Xu, Tony, et al.
Published: (2026)
by: Xu, Tony, et al.
Published: (2026)
Stabilizing Decentralized Federated Fine-Tuning via Topology-Aware Alternating LoRA
by: Wang, Xiaoyu, et al.
Published: (2026)
by: Wang, Xiaoyu, et al.
Published: (2026)
Faster Distributed Inference-Only Recommender Systems via Bounded Lag Synchronous Collectives
by: Dichev, Kiril, et al.
Published: (2025)
by: Dichev, Kiril, et al.
Published: (2025)
Flame: Simplifying Topology Extension in Federated Learning
by: Daga, Harshit, et al.
Published: (2023)
by: Daga, Harshit, et al.
Published: (2023)
Low-Communication Resilient Distributed Estimation Algorithm Based on Memory Mechanism
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Topology-Aware Knowledge Propagation in Decentralized Learning
by: Sakarvadia, Mansi, et al.
Published: (2025)
by: Sakarvadia, Mansi, et al.
Published: (2025)
Minder: Faulty Machine Detection for Large-scale Distributed Model Training
by: Deng, Yangtao, et al.
Published: (2024)
by: Deng, Yangtao, et al.
Published: (2024)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Dynamic Topology Optimization for Non-IID Data in Decentralized Learning
by: Cox, Bart, et al.
Published: (2026)
by: Cox, Bart, et al.
Published: (2026)
Approximate Agreement Algorithms for Byzantine Collaborative Learning
by: Cambus, Mélanie, et al.
Published: (2025)
by: Cambus, Mélanie, et al.
Published: (2025)
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
by: Li, Shigang, et al.
Published: (2019)
by: Li, Shigang, et al.
Published: (2019)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
by: Gupta, Vima, et al.
Published: (2024)
by: Gupta, Vima, et al.
Published: (2024)
Overcoming Challenges of Partial Client Participation in Federated Learning : A Comprehensive Review
by: Sen, Mrinmay, et al.
Published: (2025)
by: Sen, Mrinmay, et al.
Published: (2025)
Topology-aware Federated Learning in Edge Computing: A Comprehensive Survey
by: Wu, Jiajun, et al.
Published: (2023)
by: Wu, Jiajun, et al.
Published: (2023)
Beyond the Federation: Topology-aware Federated Learning for Generalization to Unseen Clients
by: Ma, Mengmeng, et al.
Published: (2024)
by: Ma, Mengmeng, et al.
Published: (2024)
Incentivizing Permissionless Distributed Learning of LLMs
by: Lidin, Joel, et al.
Published: (2025)
by: Lidin, Joel, et al.
Published: (2025)
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference
by: Meng, Han, et al.
Published: (2026)
by: Meng, Han, et al.
Published: (2026)
An Efficient Subspace Algorithm for Federated Learning on Heterogeneous Data
by: Zhang, Jiaojiao, et al.
Published: (2025)
by: Zhang, Jiaojiao, et al.
Published: (2025)
Machine Learning for Consistency Violation Faults Analysis
by: Giri, Kamal, et al.
Published: (2025)
by: Giri, Kamal, et al.
Published: (2025)
Proof-of-Collaborative-Learning: A Multi-winner Federated Learning Consensus Algorithm
by: Sokhankhosh, Amirreza, et al.
Published: (2024)
by: Sokhankhosh, Amirreza, et al.
Published: (2024)
Similar Items
-
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021) -
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
by: Bambhaniya, Abhimanyu, et al.
Published: (2024) -
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving
by: Wu, Hanjiang, et al.
Published: (2026) -
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023) -
MIST: A Co-Design Framework for Heterogeneous, Multi-Stage LLM Inference
by: Bambhaniya, Abhimanyu Rajeshkumar, et al.
Published: (2025)