Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
Fuente:
arXiv
Saved in:
| Main Authors: | Feng, Wenjiao, Xiao, Rongxing, Li, Zonghang, Yu, Hongfang, Sun, Gang, Luo, Long, Guizani, Mohsen, Ho, Qirong, Liu, Steve |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
by: Li, Zonghang, et al.
Published: (2025)
by: Li, Zonghang, et al.
Published: (2025)
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
by: Schroeder, Christian, et al.
Published: (2023)
by: Schroeder, Christian, et al.
Published: (2023)
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
by: Zhao, Mingjie, et al.
Published: (2026)
by: Zhao, Mingjie, et al.
Published: (2026)
Serial Parallel Reliability Redundancy Allocation Optimization for Energy Efficient and Fault Tolerant Cloud Computing
by: Krishna, Gutha Jaya
Published: (2024)
by: Krishna, Gutha Jaya
Published: (2024)
Multistep schemes for solving backward stochastic differential equations on GPU
by: Kapllani, Lorenc, et al.
Published: (2019)
by: Kapllani, Lorenc, et al.
Published: (2019)
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training
by: Sakip, Akhmed, et al.
Published: (2026)
by: Sakip, Akhmed, et al.
Published: (2026)
RapidStream IR: Infrastructure for FPGA High-Level Physical Synthesis
by: Lau, Jason, et al.
Published: (2024)
by: Lau, Jason, et al.
Published: (2024)
Training Diffusion Models with Federated Learning
by: de Goede, Matthijs, et al.
Published: (2024)
by: de Goede, Matthijs, et al.
Published: (2024)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
by: Xin, Jihao, et al.
Published: (2023)
by: Xin, Jihao, et al.
Published: (2023)
Asynchronous Multi-Server Federated Learning for Geo-Distributed Clients
by: Zuo, Yuncong, et al.
Published: (2024)
by: Zuo, Yuncong, et al.
Published: (2024)
The High Cost of Keeping Warm: Characterizing Overhead in Serverless Autoscaling Policies
by: Kondrashov, Leonid, et al.
Published: (2025)
by: Kondrashov, Leonid, et al.
Published: (2025)
EPARA: Parallelizing Categorized AI Inference in Edge Clouds
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
SPARK: Igniting Communication-Efficient Decentralized Learning via Stage-wise Projected NTK and Accelerated Regularization
by: Xia, Li
Published: (2025)
by: Xia, Li
Published: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
by: Li, Rongzhi, et al.
Published: (2025)
by: Li, Rongzhi, et al.
Published: (2025)
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
by: Chen, Jinfan, et al.
Published: (2023)
by: Chen, Jinfan, et al.
Published: (2023)
Self-adaptive, Requirements-driven Autoscaling of Microservices
by: Nunes, João Paulo Karol Santos, et al.
Published: (2024)
by: Nunes, João Paulo Karol Santos, et al.
Published: (2024)
How Machine Learning-Data Driven Replication Strategies Enhance Fault Tolerance in Large-Scale Distributed Systems
by: Murimi, Almond Kiruthu
Published: (2025)
by: Murimi, Almond Kiruthu
Published: (2025)
Daedalus: Self-Adaptive Horizontal Autoscaling for Resource Efficiency of Distributed Stream Processing Systems
by: Pfister, Benjamin J. J., et al.
Published: (2024)
by: Pfister, Benjamin J. J., et al.
Published: (2024)
Securing Federated Sensitive Topic Classification against Poisoning Attacks
by: Chu, Tianyue, et al.
Published: (2022)
by: Chu, Tianyue, et al.
Published: (2022)
Bridging Generalization Gap of Heterogeneous Federated Clients Using Generative Models
by: Niu, Ziru, et al.
Published: (2025)
by: Niu, Ziru, et al.
Published: (2025)
Aergia: Leveraging Heterogeneity in Federated Learning Systems
by: Cox, Bart, et al.
Published: (2022)
by: Cox, Bart, et al.
Published: (2022)
Roadmap for Edge AI: A Dagstuhl Perspective
by: Ding, Aaron Yi, et al.
Published: (2021)
by: Ding, Aaron Yi, et al.
Published: (2021)
Towards Optimal Heterogeneous Client Sampling in Multi-Model Federated Learning
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
Parameterizing Federated Continual Learning for Reproducible Research
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Asynchronous Byzantine Federated Learning
by: Cox, Bart, et al.
Published: (2024)
by: Cox, Bart, et al.
Published: (2024)
Hyper-parameter Optimization for Federated Learning with Step-wise Adaptive Mechanism
by: Saadati, Yasaman, et al.
Published: (2024)
by: Saadati, Yasaman, et al.
Published: (2024)
A Domain-Driven Design Simulator for Business Logic-Rich Microservice Systems
by: Pereira, Daniel da Palma, et al.
Published: (2026)
by: Pereira, Daniel da Palma, et al.
Published: (2026)
Connecting Large Language Model Agent to High Performance Computing Resource
by: Ma, Heng, et al.
Published: (2025)
by: Ma, Heng, et al.
Published: (2025)
Streaming REST APIs for Large Financial Transaction Exports from Relational Databases
by: Kandiraju, Abhiram
Published: (2026)
by: Kandiraju, Abhiram
Published: (2026)
Partitioned Neural Network Training via Synthetic Intermediate Labels
by: Karadağ, Cevat Volkan, et al.
Published: (2024)
by: Karadağ, Cevat Volkan, et al.
Published: (2024)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
by: Wang, Zheng, et al.
Published: (2025)
by: Wang, Zheng, et al.
Published: (2025)
Uncertainty Estimation in Multi-Agent Distributed Learning for AI-Enabled Edge Devices
by: Radchenko, Gleb, et al.
Published: (2024)
by: Radchenko, Gleb, et al.
Published: (2024)
FedPLT: Scalable, Resource-Efficient, and Heterogeneity-Aware Federated Learning via Partial Layer Training
by: Dabaja, Ahmad, et al.
Published: (2026)
by: Dabaja, Ahmad, et al.
Published: (2026)
HGraphScale: Hierarchical Graph Learning for Autoscaling Microservice Applications in Container-based Cloud Computing
by: Fang, Zhengxin, et al.
Published: (2025)
by: Fang, Zhengxin, et al.
Published: (2025)
CodeCRDT: Observation-Driven Coordination for Multi-Agent LLM Code Generation
by: Pugachev, Sergey
Published: (2025)
by: Pugachev, Sergey
Published: (2025)
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
by: Parmar, Abhinav Singh
Published: (2026)
by: Parmar, Abhinav Singh
Published: (2026)
nvidia-pcm: A D-Bus-Driven Platform Configuration Manager for OpenBMC Environments
by: Singh, Harinder
Published: (2026)
by: Singh, Harinder
Published: (2026)
SparkAttention: High-Performance Multi-Head Attention for Large Models on Volta GPU Architecture
by: Xu, Youxuan, et al.
Published: (2025)
by: Xu, Youxuan, et al.
Published: (2025)
Similar Items
-
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024) -
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024) -
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
by: Li, Zonghang, et al.
Published: (2025) -
Comparison of Autoscaling Frameworks for Containerised Machine-Learning-Applications in a Local and Cloud Environment
by: Schroeder, Christian, et al.
Published: (2023) -
SI-ChainFL: Shapley-Incentivized Secure Federated Learning for High-Speed Rail Data Sharing
by: Zhao, Mingjie, et al.
Published: (2026)