History Rhymes: Accelerating LLM Reinforcement Learning with RhymeRL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Jingkai, Li, Tianjian, Feng, Erhu, Du, Dong, Liu, Qian, Liu, Tao, Xia, Yubin, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
von: Feng, Dahu, et al.
Veröffentlicht: (2025)
von: Feng, Dahu, et al.
Veröffentlicht: (2025)
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025)
von: Chen, Le, et al.
Veröffentlicht: (2025)
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
von: Liu, Qingyuan, et al.
Veröffentlicht: (2024)
von: Liu, Qingyuan, et al.
Veröffentlicht: (2024)
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
von: He, Yongjun, et al.
Veröffentlicht: (2025)
von: He, Yongjun, et al.
Veröffentlicht: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
von: Li, Pengbo, et al.
Veröffentlicht: (2026)
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing
von: Zhu, Jianyong, et al.
Veröffentlicht: (2026)
von: Zhu, Jianyong, et al.
Veröffentlicht: (2026)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
von: Wang, Zhixin, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Accelerating Compound LLM Training Workloads with Maestro
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
von: Yuan, Xiulong, et al.
Veröffentlicht: (2026)
PICO: Accelerating All k-Core Paradigms on GPU
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
PROSERVE: Unified Multi-Priority Request Scheduling for LLM Serving
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
von: Huang, Weizhe, et al.
Veröffentlicht: (2025)
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Polar: Agentic RL on Any Harness at Scale
von: Xu, Binfeng, et al.
Veröffentlicht: (2026)
von: Xu, Binfeng, et al.
Veröffentlicht: (2026)
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Minions: Accelerating Large Language Model Inference with Aggregated Speculative Execution
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
von: Wang, Siqi, et al.
Veröffentlicht: (2024)
Xorbits: Automating Operator Tiling for Distributed Data Science
von: Lu, Weizheng, et al.
Veröffentlicht: (2023)
von: Lu, Weizheng, et al.
Veröffentlicht: (2023)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
von: Wang, Jun, et al.
Veröffentlicht: (2025)
von: Wang, Jun, et al.
Veröffentlicht: (2025)
A Preliminary Study on Accelerating Simulation Optimization with GPU Implementation
von: He, Jinghai, et al.
Veröffentlicht: (2024)
von: He, Jinghai, et al.
Veröffentlicht: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
Hyperion: Hierarchical Scheduling for Parallel LLM Acceleration in Multi-tier Networks
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
von: Ma, Mulei, et al.
Veröffentlicht: (2025)
Parallel Collaborative ADMM Privacy Computing and Adaptive GPU Acceleration for Distributed Edge Networks
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
von: Xia, Mengchun, et al.
Veröffentlicht: (2026)
Towards Lock Modularization for Heterogeneous Environments
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
von: Zhang, Hanze, et al.
Veröffentlicht: (2025)
HLoRA: Efficient Federated Learning System for LLM Heterogeneous Fine-Tuning
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
von: Liu, Qianli, et al.
Veröffentlicht: (2025)
LLM-Enhanced Deep Reinforcement Learning for Task Offloading in Collaborative Edge Computing
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
WWW.Serve: Interconnecting Global LLM Services through Decentralization
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
von: Wang, Huanyu, et al.
Veröffentlicht: (2026)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
LAAFD: LLM-based Agents for Accelerated FPGA Design
von: Moraru, Maxim, et al.
Veröffentlicht: (2026)
von: Moraru, Maxim, et al.
Veröffentlicht: (2026)
PlexRL: Cluster-Level Orchestration of Serviceized LLM Execution for RLVR
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
AcceLLM: Accelerating LLM Inference using Redundancy for Load Balancing and Data Locality
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
von: Bournias, Ilias, et al.
Veröffentlicht: (2024)
A Reinforcement Learning Based Backfilling Strategy for HPC Batch Jobs
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
von: Kolker-Hicks, Elliot, et al.
Veröffentlicht: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
von: Zhao, Bohan, et al.
Veröffentlicht: (2025)
Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
von: Liu, Zhibang, et al.
Veröffentlicht: (2025)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
exa-AMD: A Scalable Workflow for Accelerating AI-Assisted Materials Discovery and Design
von: Moraru, Maxim, et al.
Veröffentlicht: (2025)
von: Moraru, Maxim, et al.
Veröffentlicht: (2025)
GPU-Accelerated Batch-Dynamic Subgraph Matching
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
von: Qiu, Linshan, et al.
Veröffentlicht: (2024)
Achieving Dimension-Free Communication in Federated Learning via Zeroth-Order Optimization
von: Li, Zhe, et al.
Veröffentlicht: (2024)
von: Li, Zhe, et al.
Veröffentlicht: (2024)
UniPar: A Unified LLM-Based Framework for Parallel and Accelerated Code Translation in HPC
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
von: Bitan, Tomer, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
von: Feng, Dahu, et al.
Veröffentlicht: (2025) -
Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference
von: Chen, Le, et al.
Veröffentlicht: (2025) -
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
von: Liu, Qingyuan, et al.
Veröffentlicht: (2024) -
HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments
von: He, Yongjun, et al.
Veröffentlicht: (2025) -
Schedule-Level Shared-Prefix Reuse for LLM RL Training
von: Li, Pengbo, et al.
Veröffentlicht: (2026)