Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Long, Wang, Qinghe, Zhu, Jiaan, Bai, Youhui, Jin, Zewen, Ruan, Chaoyi, Wang, Shengnan, Li, Cheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
por: Xu, Guanbin, et al.
Publicado: (2026)
por: Xu, Guanbin, et al.
Publicado: (2026)
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
por: Ruan, Chaoyi, et al.
Publicado: (2026)
por: Ruan, Chaoyi, et al.
Publicado: (2026)
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
por: Ai, Xin, et al.
Publicado: (2024)
por: Ai, Xin, et al.
Publicado: (2024)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
por: Wang, Haiquan, et al.
Publicado: (2024)
por: Wang, Haiquan, et al.
Publicado: (2024)
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
por: Wang, Zhigang, et al.
Publicado: (2024)
por: Wang, Zhigang, et al.
Publicado: (2024)
DreamDDP: Accelerating Data Parallel Distributed LLM Training with Layer-wise Scheduled Partial Synchronization
por: Tang, Zhenheng, et al.
Publicado: (2025)
por: Tang, Zhenheng, et al.
Publicado: (2025)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
por: Gao, Wei, et al.
Publicado: (2025)
por: Gao, Wei, et al.
Publicado: (2025)
LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
por: Gu, Diandian, et al.
Publicado: (2024)
por: Gu, Diandian, et al.
Publicado: (2024)
Synergistic Tensor and Pipeline Parallelism
por: Qi, Mengshi, et al.
Publicado: (2025)
por: Qi, Mengshi, et al.
Publicado: (2025)
ElasWave: An Elastic-Native System for Scalable Hybrid-Parallel Training
por: Kang, Xueze, et al.
Publicado: (2025)
por: Kang, Xueze, et al.
Publicado: (2025)
Reaching Agreement Among Reasoning LLM Agents
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
por: Chen, Haoyu, et al.
Publicado: (2025)
por: Chen, Haoyu, et al.
Publicado: (2025)
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
por: Srivatsa, Vikranth, et al.
Publicado: (2026)
por: Srivatsa, Vikranth, et al.
Publicado: (2026)
SPPO:Efficient Long-sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading
por: Chen, Qiaoling, et al.
Publicado: (2025)
por: Chen, Qiaoling, et al.
Publicado: (2025)
ZeroPP: Unleashing Exceptional Parallelism Efficiency through Tensor-Parallelism-Free Methodology
por: Tang, Ding, et al.
Publicado: (2024)
por: Tang, Ding, et al.
Publicado: (2024)
InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
por: Chen, Qiaoling, et al.
Publicado: (2024)
por: Chen, Qiaoling, et al.
Publicado: (2024)
Accelerating Compound LLM Training Workloads with Maestro
por: Yuan, Xiulong, et al.
Publicado: (2026)
por: Yuan, Xiulong, et al.
Publicado: (2026)
AnchorTP: Resilient LLM Inference with State-Preserving Elastic Tensor Parallelism
por: Xu, Wendong, et al.
Publicado: (2025)
por: Xu, Wendong, et al.
Publicado: (2025)
Accelerating Distributed MoE Training and Inference with Lina
por: Li, Jiamin, et al.
Publicado: (2022)
por: Li, Jiamin, et al.
Publicado: (2022)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
por: Xia, Mengchun, et al.
Publicado: (2026)
por: Xia, Mengchun, et al.
Publicado: (2026)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
por: Li, Shengwei, et al.
Publicado: (2023)
por: Li, Shengwei, et al.
Publicado: (2023)
Federated Learning Using Coupled Tensor Train Decomposition
por: Zhang, Xiangtao, et al.
Publicado: (2024)
por: Zhang, Xiangtao, et al.
Publicado: (2024)
DynaServe: Unified and Elastic Execution for Dynamic Disaggregated LLM Serving
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
por: Wang, Yujie, et al.
Publicado: (2024)
por: Wang, Yujie, et al.
Publicado: (2024)
StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training
por: Liu, Ziming, et al.
Publicado: (2024)
por: Liu, Ziming, et al.
Publicado: (2024)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
por: Liu, Man, et al.
Publicado: (2026)
por: Liu, Man, et al.
Publicado: (2026)
HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
por: Liang, Antian, et al.
Publicado: (2025)
por: Liang, Antian, et al.
Publicado: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
por: Li, Haoyang, et al.
Publicado: (2024)
por: Li, Haoyang, et al.
Publicado: (2024)
HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism
por: Zhang, Geng, et al.
Publicado: (2025)
por: Zhang, Geng, et al.
Publicado: (2025)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
por: Wu, Tianyuan, et al.
Publicado: (2025)
por: Wu, Tianyuan, et al.
Publicado: (2025)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
por: Yang, Yi, et al.
Publicado: (2025)
por: Yang, Yi, et al.
Publicado: (2025)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
por: Chen, Chang, et al.
Publicado: (2025)
por: Chen, Chang, et al.
Publicado: (2025)
InfiniPipe: Elastic Pipeline Parallelism for Efficient Variable-Length Long-Context LLM Training
por: Wang, Shiju, et al.
Publicado: (2025)
por: Wang, Shiju, et al.
Publicado: (2025)
Occult: Optimizing Collaborative Communication across Experts for Accelerated Parallel MoE Training and Inference
por: Luo, Shuqing, et al.
Publicado: (2025)
por: Luo, Shuqing, et al.
Publicado: (2025)
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
por: Wang, Hongyu, et al.
Publicado: (2026)
por: Wang, Hongyu, et al.
Publicado: (2026)
Minimizing Communication for Parallel Symmetric Tensor Times Same Vector Computation
por: Daas, Hussam Al, et al.
Publicado: (2025)
por: Daas, Hussam Al, et al.
Publicado: (2025)
Revisiting Parameter Server in LLM Post-Training
por: Wan, Xinyi, et al.
Publicado: (2026)
por: Wan, Xinyi, et al.
Publicado: (2026)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
Accelerating Sparse MTTKRP for Small Tensor Decomposition on GPU
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
por: Wijeratne, Sasindu, et al.
Publicado: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
por: Liu, Zhibang, et al.
Publicado: (2025)
por: Liu, Zhibang, et al.
Publicado: (2025)
Ejemplares similares
-
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
por: Xu, Guanbin, et al.
Publicado: (2026) -
RL over Commodity Networks: Overcoming the Bandwidth Barrier with Lossless Sparse Deltas
por: Ruan, Chaoyi, et al.
Publicado: (2026) -
NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor Parallelism
por: Ai, Xin, et al.
Publicado: (2024) -
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
por: Wang, Haiquan, et al.
Publicado: (2024) -
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
por: Wang, Zhigang, et al.
Publicado: (2024)