QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Hao, Wang, Hao, Liu, Jiacheng, Li, Lujun, Zhu, Qiyuan, Liu, Bei, Xu, Binxing, Wang, Lei, Yang, Xintong, Lin, Sida, Han, Sirui, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
von: Yang, Xintong, et al.
Veröffentlicht: (2026)
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
von: Xu, Binxing, et al.
Veröffentlicht: (2026)
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025)
von: Li, Lujun, et al.
Veröffentlicht: (2025)
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
von: Zhang, Hao, et al.
Veröffentlicht: (2026)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
von: Gao, Wei, et al.
Veröffentlicht: (2025)
von: Gao, Wei, et al.
Veröffentlicht: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Towards Better RL Training Data Utilization via Second-Order Rollout
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
Patch the Distribution Mismatch: RL Rewriting Agent for Stable Off-Policy SFT
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
von: Wang, Jiacheng, et al.
Veröffentlicht: (2026)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Unleashing Efficient Asynchronous RL Post-Training via Staleness-Constrained Rollout Coordination
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
von: Li, Haoyang, et al.
Veröffentlicht: (2026)
Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decoding
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
von: Iso, Hayate, et al.
Veröffentlicht: (2026)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
von: Gai, Jiading, et al.
Veröffentlicht: (2026)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
von: Liu, Bingshuai, et al.
Veröffentlicht: (2025)
$V_{0.5}$: Generalist Value Model as a Prior for Sparse RL Rollouts
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
von: Zhang, Yi-Kai, et al.
Veröffentlicht: (2026)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
von: Xi, Haocheng, et al.
Veröffentlicht: (2026)
EchoRL: Reinforcement Learning via Rollout Echoing
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
von: Bi, Jinhe, et al.
Veröffentlicht: (2026)
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
von: Wang, Liangyu, et al.
Veröffentlicht: (2025)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
von: Zhang, Yiqi, et al.
Veröffentlicht: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
MindSpeed RL: Distributed Dataflow for Scalable and Efficient RL Training on Ascend NPU Cluster
von: Feng, Laingjun, et al.
Veröffentlicht: (2025)
von: Feng, Laingjun, et al.
Veröffentlicht: (2025)
Domain-Aware RAG: MoL-Enhanced RL for Efficient Training and Scalable Retrieval
von: Lin, Hao, et al.
Veröffentlicht: (2025)
von: Lin, Hao, et al.
Veröffentlicht: (2025)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
von: Zhang, Zili, et al.
Veröffentlicht: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
von: Luo, Sijia, et al.
Veröffentlicht: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
Beat the long tail: Distribution-Aware Speculative Decoding for RL Training
von: Shao, Zelei, et al.
Veröffentlicht: (2025)
von: Shao, Zelei, et al.
Veröffentlicht: (2025)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
von: Gao, Wei, et al.
Veröffentlicht: (2026)
von: Gao, Wei, et al.
Veröffentlicht: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
von: Chen, Wanyi, et al.
Veröffentlicht: (2026)
GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
von: Hu, Yuelin, et al.
Veröffentlicht: (2026)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
von: Liu, Jie, et al.
Veröffentlicht: (2025)
von: Liu, Jie, et al.
Veröffentlicht: (2025)
SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
von: Cao, Shiyi, et al.
Veröffentlicht: (2025)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
DiffusionRL: Efficient Training of Diffusion Policies for Robotic Grasping Using RL-Adapted Large-Scale Datasets
von: Makarova, Maria, et al.
Veröffentlicht: (2025)
von: Makarova, Maria, et al.
Veröffentlicht: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
s3: You Don't Need That Much Data to Train a Search Agent via RL
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2025)
Delta Decompression for MoE-based LLMs Compression
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
von: Yang, Xintong, et al.
Veröffentlicht: (2026) -
Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
von: Xu, Binxing, et al.
Veröffentlicht: (2026) -
BTC-LLM: Efficient Sub-1-Bit LLM Quantization via Learnable Transformation and Binary Codebook
von: Gu, Hao, et al.
Veröffentlicht: (2025) -
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
von: Li, Lujun, et al.
Veröffentlicht: (2025) -
ProRL Agent: Rollout-as-a-Service for RL Training of Multi-Turn LLM Agents
von: Zhang, Hao, et al.
Veröffentlicht: (2026)