A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaocan, Wu, Shiliang, Shen, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
by: Jung, Hyunji, et al.
Published: (2026)
by: Jung, Hyunji, et al.
Published: (2026)
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024)
by: Ma, Jeffrey, et al.
Published: (2024)
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025)
by: Sheng, Guangming, et al.
Published: (2025)
Ravnest: Decentralized Asynchronous Training on Heterogeneous Devices
by: Menon, Anirudh Rajiv, et al.
Published: (2024)
by: Menon, Anirudh Rajiv, et al.
Published: (2024)
FedFa: A Fully Asynchronous Training Paradigm for Federated Learning
by: Xu, Haotian, et al.
Published: (2024)
by: Xu, Haotian, et al.
Published: (2024)
SEAFL: Enhancing Efficiency in Semi-Asynchronous Federated Learning through Adaptive Aggregation and Selective Training
by: Islam, Md Sirajul, et al.
Published: (2025)
by: Islam, Md Sirajul, et al.
Published: (2025)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
BitPipe: Bidirectional Interleaved Pipeline Parallelism for Accelerating Large Models Training
by: Wu, Houming, et al.
Published: (2024)
by: Wu, Houming, et al.
Published: (2024)
MSPipe: Efficient Temporal GNN Training via Staleness-Aware Pipeline
by: Sheng, Guangming, et al.
Published: (2024)
by: Sheng, Guangming, et al.
Published: (2024)
TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models Training
by: Wu, Houming, et al.
Published: (2025)
by: Wu, Houming, et al.
Published: (2025)
Robust LLM Training Infrastructure at ByteDance
by: Wan, Borui, et al.
Published: (2025)
by: Wan, Borui, et al.
Published: (2025)
FedPBS: Proximal-Balanced Scaling Federated Learning Model for Robust Personalized Training for Non-IID Data
by: AbouNassar, Eman M., et al.
Published: (2026)
by: AbouNassar, Eman M., et al.
Published: (2026)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
Efficient Asynchronous Federated Learning with Sparsification and Quantization
by: Jia, Juncheng, et al.
Published: (2023)
by: Jia, Juncheng, et al.
Published: (2023)
GPRat: Gaussian Process Regression with Asynchronous Tasks
by: Helmann, Maksim, et al.
Published: (2025)
by: Helmann, Maksim, et al.
Published: (2025)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
by: Chen, Yanxi, et al.
Published: (2023)
by: Chen, Yanxi, et al.
Published: (2023)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
by: Chen, Le, et al.
Published: (2025)
by: Chen, Le, et al.
Published: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
Mitigating Persistent Client Dropout in Asynchronous Decentralized Federated Learning
by: Stępka, Ignacy, et al.
Published: (2025)
by: Stępka, Ignacy, et al.
Published: (2025)
Role-Based Fault Tolerance System for LLM RL Post-Training
by: Chen, Zhenqian, et al.
Published: (2025)
by: Chen, Zhenqian, et al.
Published: (2025)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
by: Zhou, Jiecheng, et al.
Published: (2025)
by: Zhou, Jiecheng, et al.
Published: (2025)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
HASFL: Heterogeneity-aware Split Federated Learning over Edge Computing Systems
by: Lin, Zheng, et al.
Published: (2025)
by: Lin, Zheng, et al.
Published: (2025)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
Empirical Analysis of Asynchronous Federated Learning on Heterogeneous Devices: Efficiency, Fairness, and Privacy Trade-offs
by: Mohammadi, Samaneh, et al.
Published: (2025)
by: Mohammadi, Samaneh, et al.
Published: (2025)
Speeding up Policy Simulation in Supply Chain RL
by: Farias, Vivek, et al.
Published: (2024)
by: Farias, Vivek, et al.
Published: (2024)
Mind the Gap: Revealing Inconsistencies Across Heterogeneous AI Accelerators
by: Wen, Elliott, et al.
Published: (2025)
by: Wen, Elliott, et al.
Published: (2025)
CG-FedLLM: How to Compress Gradients in Federated Fune-tuning for Large Language Models
by: Wu, Huiwen, et al.
Published: (2024)
by: Wu, Huiwen, et al.
Published: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
by: Shen, Zheyu, et al.
Published: (2025)
by: Shen, Zheyu, et al.
Published: (2025)
Efficient Training on Multiple Consumer GPUs with RoundPipe
by: Luo, Yibin, et al.
Published: (2026)
by: Luo, Yibin, et al.
Published: (2026)
Federated style aware transformer aggregation of representations
by: Jeon, Mincheol, et al.
Published: (2025)
by: Jeon, Mincheol, et al.
Published: (2025)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
by: Zhang, Jiyuan, et al.
Published: (2026)
by: Zhang, Jiyuan, et al.
Published: (2026)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
by: Yang, Hanmei, et al.
Published: (2024)
by: Yang, Hanmei, et al.
Published: (2024)
AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM Serving
by: Xu, Tianhao, et al.
Published: (2026)
by: Xu, Tianhao, et al.
Published: (2026)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
by: Ge, Hao, et al.
Published: (2025)
by: Ge, Hao, et al.
Published: (2025)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
by: Zeng, Zhichen, et al.
Published: (2026)
by: Zeng, Zhichen, et al.
Published: (2026)
Low-Rank GEMM: Efficient Matrix Multiplication via Low-Rank Approximation with FP8 Acceleration
by: Metere, Alfredo
Published: (2025)
by: Metere, Alfredo
Published: (2025)
Similar Items
-
Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis Rotation
by: Jung, Hyunji, et al.
Published: (2026) -
FedStaleWeight: Buffered Asynchronous Federated Learning with Fair Aggregation via Staleness Reweighting
by: Ma, Jeffrey, et al.
Published: (2024) -
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025) -
Ravnest: Decentralized Asynchronous Training on Heterogeneous Devices
by: Menon, Anirudh Rajiv, et al.
Published: (2024) -
FedFa: A Fully Asynchronous Training Paradigm for Federated Learning
by: Xu, Haotian, et al.
Published: (2024)