QuRL: Efficient Reinforcement Learning with Quantized Rollout
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Yuhang, Elangovan, Reena, Dong, Xin, Panda, Priyadarshini, Khailany, Brucek |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
por: Elangovan, Reena, et al.
Publicado: (2025)
por: Elangovan, Reena, et al.
Publicado: (2025)
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
por: Li, Yuhang, et al.
Publicado: (2024)
por: Li, Yuhang, et al.
Publicado: (2024)
ESPACE: Dimensionality Reduction of Activations for Model Compression
por: Sakr, Charbel, et al.
Publicado: (2024)
por: Sakr, Charbel, et al.
Publicado: (2024)
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
por: Li, Yuhang, et al.
Publicado: (2025)
por: Li, Yuhang, et al.
Publicado: (2025)
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
por: Song, Zhiye, et al.
Publicado: (2025)
por: Song, Zhiye, et al.
Publicado: (2025)
SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
por: Fan, Zichen, et al.
Publicado: (2025)
por: Fan, Zichen, et al.
Publicado: (2025)
ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning Models
por: Ramachandran, Akshat, et al.
Publicado: (2025)
por: Ramachandran, Akshat, et al.
Publicado: (2025)
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference
por: Hooper, Coleman, et al.
Publicado: (2025)
por: Hooper, Coleman, et al.
Publicado: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
por: Bi, Jinhe, et al.
Publicado: (2026)
por: Bi, Jinhe, et al.
Publicado: (2026)
Memba: Membrane-driven Parameter-Efficient Fine-Tuning for Mamba
por: Lee, Donghyun, et al.
Publicado: (2025)
por: Lee, Donghyun, et al.
Publicado: (2025)
Optimal Brain Decomposition for Accurate LLM Low-Rank Approximation
por: Li, Yuhang, et al.
Publicado: (2026)
por: Li, Yuhang, et al.
Publicado: (2026)
DuoGPT: Training-free Dual Sparsity through Activation-aware Pruning in LLMs
por: Yin, Ruokai, et al.
Publicado: (2025)
por: Yin, Ruokai, et al.
Publicado: (2025)
TurboSAT: Gradient-Guided Boolean Satisfiability Accelerated on GPU-CPU Hybrid System
por: Dai, Steve, et al.
Publicado: (2025)
por: Dai, Steve, et al.
Publicado: (2025)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
por: Liu, Bingshuai, et al.
Publicado: (2025)
por: Liu, Bingshuai, et al.
Publicado: (2025)
LLM4Cov: Execution-Aware Agentic Learning for High-coverage Testbench Generation
por: Zhang, Hejia, et al.
Publicado: (2026)
por: Zhang, Hejia, et al.
Publicado: (2026)
ReSpike: Residual Frames-based Hybrid Spiking Neural Networks for Efficient Action Recognition
por: Xiao, Shiting, et al.
Publicado: (2024)
por: Xiao, Shiting, et al.
Publicado: (2024)
QaRL: Rollout-Aligned Quantization-Aware RL for Fast and Stable Training under Training--Inference Mismatch
por: Gu, Hao, et al.
Publicado: (2026)
por: Gu, Hao, et al.
Publicado: (2026)
Selective Rollout: Mid-Trajectory Termination for Multi-Sample Agent RL
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
por: Zhai, Zhiyuan, et al.
Publicado: (2026)
RL2ML: Finite-Rollout Surrogate Objectives from Reinforcement Learning to Maximum Likelihood
por: Zheng, Yifu
Publicado: (2026)
por: Zheng, Yifu
Publicado: (2026)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
por: Liu, Rui, et al.
Publicado: (2025)
por: Liu, Rui, et al.
Publicado: (2025)
Heddle: A Distributed Orchestration System for Agentic RL Rollout
por: Zhang, Zili, et al.
Publicado: (2026)
por: Zhang, Zili, et al.
Publicado: (2026)
On Rollouts in Model-Based Reinforcement Learning
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
por: Frauenknecht, Bernd, et al.
Publicado: (2025)
ACE-RTL: When Agentic Context Evolution Meets RTL-Specialized LLMs
por: Deng, Chenhui, et al.
Publicado: (2026)
por: Deng, Chenhui, et al.
Publicado: (2026)
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
por: Liu, Wenpu, et al.
Publicado: (2026)
por: Liu, Wenpu, et al.
Publicado: (2026)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
por: Xi, Haocheng, et al.
Publicado: (2026)
por: Xi, Haocheng, et al.
Publicado: (2026)
Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts
por: Luo, Sijia, et al.
Publicado: (2026)
por: Luo, Sijia, et al.
Publicado: (2026)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
por: Jiang, Zhida, et al.
Publicado: (2026)
por: Jiang, Zhida, et al.
Publicado: (2026)
BubbleSpec: Turning Long-Tail Bubbles into Speculative Rollout Drafts for Synchronous Reinforcement Learning
por: Xu, Yuhang, et al.
Publicado: (2026)
por: Xu, Yuhang, et al.
Publicado: (2026)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
por: Xu, Yixuan Even, et al.
Publicado: (2025)
por: Xu, Yixuan Even, et al.
Publicado: (2025)
BREAD: Branched Rollouts from Expert Anchors Bridge SFT & RL for Reasoning
por: Zhang, Xuechen, et al.
Publicado: (2025)
por: Zhang, Xuechen, et al.
Publicado: (2025)
DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts
por: Gai, Jiading, et al.
Publicado: (2026)
por: Gai, Jiading, et al.
Publicado: (2026)
ImagineBench: Evaluating Reinforcement Learning with Large Language Model Rollouts
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
por: Pang, Jing-Cheng, et al.
Publicado: (2025)
When In-memory Computing Meets Spiking Neural Networks -- A Perspective on Device-Circuit-System-and-Algorithm Co-design
por: Moitra, Abhishek, et al.
Publicado: (2024)
por: Moitra, Abhishek, et al.
Publicado: (2024)
Train Less, Learn More: Adaptive Efficient Rollout Optimization for Group-Based Reinforcement Learning
por: Zhang, Zhi, et al.
Publicado: (2026)
por: Zhang, Zhi, et al.
Publicado: (2026)
ClipFormer: Key-Value Clipping of Transformers on Memristive Crossbars for Write Noise Mitigation
por: Bhattacharjee, Abhiroop, et al.
Publicado: (2024)
por: Bhattacharjee, Abhiroop, et al.
Publicado: (2024)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
por: Zhu, Tianshu, et al.
Publicado: (2026)
por: Zhu, Tianshu, et al.
Publicado: (2026)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
por: Zheng, Haizhong, et al.
Publicado: (2025)
por: Zheng, Haizhong, et al.
Publicado: (2025)
Predicting Probabilities of Error to Combine Quantization and Early Exiting: QuEE
por: Regol, Florence, et al.
Publicado: (2024)
por: Regol, Florence, et al.
Publicado: (2024)
Regularization-based Framework for Quantization-, Fault- and Variability-Aware Training
por: Biswas, Anmol, et al.
Publicado: (2025)
por: Biswas, Anmol, et al.
Publicado: (2025)
Ejemplares similares
-
LO-BCQ: Block Clustered Quantization for 4-bit (W4A4) LLM Inference
por: Elangovan, Reena, et al.
Publicado: (2025) -
TesseraQ: Ultra Low-Bit LLM Post-Training Quantization with Block Reconstruction
por: Li, Yuhang, et al.
Publicado: (2024) -
ESPACE: Dimensionality Reduction of Activations for Model Compression
por: Sakr, Charbel, et al.
Publicado: (2024) -
GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
por: Li, Yuhang, et al.
Publicado: (2025) -
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
por: Song, Zhiye, et al.
Publicado: (2025)