DYNAMIX: RL-based Adaptive Batch Size Optimization in Distributed Machine Learning Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Dai, Yuanjun, He, Keqiang, Wang, An |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collaborative Batch Size Optimization for Federated Learning
by: Geimer, Arno, et al.
Published: (2025)
by: Geimer, Arno, et al.
Published: (2025)
FLAMMABLE: A Multi-Model Federated Learning Framework with Multi-Model Engagement and Adaptive Batch Sizes
by: Lin, Shouxu, et al.
Published: (2025)
by: Lin, Shouxu, et al.
Published: (2025)
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025)
by: Lu, Kuan-Wei, et al.
Published: (2025)
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
by: Chen, Hongyao, et al.
Published: (2025)
by: Chen, Hongyao, et al.
Published: (2025)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Enabling Large Batch Size Training for DNN Models Beyond the Memory Limit While Maintaining Performance
by: Piao, XinYu, et al.
Published: (2021)
by: Piao, XinYu, et al.
Published: (2021)
Optimal Batch Allocation for Wireless Federated Learning
by: Song, Jaeyoung, et al.
Published: (2024)
by: Song, Jaeyoung, et al.
Published: (2024)
CDFGNN: a Systematic Design of Cache-based Distributed Full-Batch Graph Neural Network Training with Communication Reduction
by: Zhang, Shuai, et al.
Published: (2024)
by: Zhang, Shuai, et al.
Published: (2024)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
by: He, Jingkai, et al.
Published: (2025)
by: He, Jingkai, et al.
Published: (2025)
Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of Things
by: Wang, Ziheng, et al.
Published: (2024)
by: Wang, Ziheng, et al.
Published: (2024)
Learning to Shard: RL for Co-optimizing the Parallelism Degrees and Per-operator Sharding Dimensions in Distributed LLM Inference
by: Yin, Ruokai, et al.
Published: (2025)
by: Yin, Ruokai, et al.
Published: (2025)
FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
by: Zhou, Liuzhi, et al.
Published: (2024)
by: Zhou, Liuzhi, et al.
Published: (2024)
Machine-Learning-Driven Runtime Optimization of BLAS Level 3 on Modern Multi-Core Systems
by: Xia, Yufan, et al.
Published: (2024)
by: Xia, Yufan, et al.
Published: (2024)
Accelerating Wireless Distributed Learning via Hybrid Split and Federated Learning Optimization
by: Guo, Kun, et al.
Published: (2025)
by: Guo, Kun, et al.
Published: (2025)
Reinforcement Learning for Adaptive Resource Scheduling in Complex System Environments
by: Li, Pochun, et al.
Published: (2024)
by: Li, Pochun, et al.
Published: (2024)
Canzona: A Unified, Asynchronous, and Load-Balanced Framework for Distributed Matrix-based Optimizers
by: Wang, Liangyu, et al.
Published: (2026)
by: Wang, Liangyu, et al.
Published: (2026)
EMO: Edge Model Overlays to Scale Model Size in Federated Learning
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-Batch
by: Bajaj, Saurabh, et al.
Published: (2024)
by: Bajaj, Saurabh, et al.
Published: (2024)
AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT Systems
by: Jia, Chentao, et al.
Published: (2023)
by: Jia, Chentao, et al.
Published: (2023)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
DynLP: Parallel Dynamic Batch Update for Label Propagation in Semi-Supervised Learning
by: Shovan, S M, et al.
Published: (2026)
by: Shovan, S M, et al.
Published: (2026)
The Streaming Batch Model for Efficient and Fault-Tolerant Heterogeneous Execution
by: Luan, Frank Sifei, et al.
Published: (2025)
by: Luan, Frank Sifei, et al.
Published: (2025)
Personalized Federated Domain-Incremental Learning based on Adaptive Knowledge Matching
by: Li, Yichen, et al.
Published: (2024)
by: Li, Yichen, et al.
Published: (2024)
Efficient Distributed Learning over Decentralized Networks with Convoluted Support Vector Machine
by: Chen, Canyi, et al.
Published: (2025)
by: Chen, Canyi, et al.
Published: (2025)
Provenance Tracking in Large-Scale Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
MergeSFL: Split Federated Learning with Feature Merging and Batch Size Regulation
by: Liao, Yunming, et al.
Published: (2023)
by: Liao, Yunming, et al.
Published: (2023)
Towards Communication-efficient Federated Learning via Sparse and Aligned Adaptive Optimization
by: Deng, Xiumei, et al.
Published: (2024)
by: Deng, Xiumei, et al.
Published: (2024)
StreamRL: Scalable, Heterogeneous, and Elastic RL for LLMs with Disaggregated Stream Generation
by: Zhong, Yinmin, et al.
Published: (2025)
by: Zhong, Yinmin, et al.
Published: (2025)
FedSAE: A Novel Self-Adaptive Federated Learning Framework in Heterogeneous Systems
by: Li, Li, et al.
Published: (2021)
by: Li, Li, et al.
Published: (2021)
Optimizing Performance on Trinity Utilizing Machine Learning, Proxy Applications and Scheduling Priorities
by: Romero, Phil
Published: (2024)
by: Romero, Phil
Published: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
Resource-Adaptive Successive Doubling for Hyperparameter Optimization with Large Datasets on High-Performance Computing Systems
by: Aach, Marcel, et al.
Published: (2024)
by: Aach, Marcel, et al.
Published: (2024)
ReSpec: Towards Optimizing Speculative Decoding in Reinforcement Learning Systems
by: Chen, Qiaoling, et al.
Published: (2025)
by: Chen, Qiaoling, et al.
Published: (2025)
FedFetch: Faster Federated Learning with Adaptive Downstream Prefetching
by: Yan, Qifan, et al.
Published: (2025)
by: Yan, Qifan, et al.
Published: (2025)
CoBo: Collaborative Learning via Bilevel Optimization
by: Hashemi, Diba, et al.
Published: (2024)
by: Hashemi, Diba, et al.
Published: (2024)
Reinforcement Learning-based Adaptive Mitigation of Uncorrected DRAM Errors in the Field
by: Boixaderas, Isaac, et al.
Published: (2024)
by: Boixaderas, Isaac, et al.
Published: (2024)
Virtual Nodes Can Help: Tackling Distribution Shifts in Federated Graph Learning
by: Fu, Xingbo, et al.
Published: (2024)
by: Fu, Xingbo, et al.
Published: (2024)
SplitFT: An Adaptive Federated Split Learning System For LLMs Fine-Tuning
by: Shan, Yimeng, et al.
Published: (2026)
by: Shan, Yimeng, et al.
Published: (2026)
Similar Items
-
Collaborative Batch Size Optimization for Federated Learning
by: Geimer, Arno, et al.
Published: (2025) -
FLAMMABLE: A Multi-Model Federated Learning Framework with Multi-Model Engagement and Adaptive Batch Sizes
by: Lin, Shouxu, et al.
Published: (2025) -
Hybrid Dual-Batch and Cyclic Progressive Learning for Efficient Distributed Training
by: Lu, Kuan-Wei, et al.
Published: (2025) -
Hybrid Batch Normalisation: Resolving the Dilemma of Batch Normalisation in Federated Learning
by: Chen, Hongyao, et al.
Published: (2025) -
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)