Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Fuente:
arXiv
Saved in:
| Main Authors: | Zheng, Haizhong, Zhou, Yang, Bartoldson, Brian R., Kailkhura, Bhavya, Lai, Fan, Zhao, Jiawei, Chen, Beidi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
by: Zheng, Haizhong, et al.
Published: (2024)
by: Zheng, Haizhong, et al.
Published: (2024)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
by: Bartoldson, Brian R., et al.
Published: (2024)
by: Bartoldson, Brian R., et al.
Published: (2024)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025)
by: Wang, Zeyu, et al.
Published: (2025)
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
by: McDonald, Tavish, et al.
Published: (2025)
by: McDonald, Tavish, et al.
Published: (2025)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
by: Cai, Zikui, et al.
Published: (2025)
by: Cai, Zikui, et al.
Published: (2025)
Jackpot: Optimal Budgeted Rejection Sampling for Extreme Actor-Policy Mismatch Reinforcement Learning
by: Chen, Zhuoming, et al.
Published: (2026)
by: Chen, Zhuoming, et al.
Published: (2026)
Leveraging Hierarchical Feature Sharing for Efficient Dataset Condensation
by: Zheng, Haizhong, et al.
Published: (2023)
by: Zheng, Haizhong, et al.
Published: (2023)
Learn To be Efficient: Build Structured Sparsity in Large Language Models
by: Zheng, Haizhong, et al.
Published: (2024)
by: Zheng, Haizhong, et al.
Published: (2024)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
by: Wang, Zijun, et al.
Published: (2025)
by: Wang, Zijun, et al.
Published: (2025)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
by: Christopher, Jacob K, et al.
Published: (2024)
by: Christopher, Jacob K, et al.
Published: (2024)
Certifiably-Robust Federated Adversarial Learning via Randomized Smoothing
by: Chen, Cheng, et al.
Published: (2021)
by: Chen, Cheng, et al.
Published: (2021)
FedCluster: Boosting the Convergence of Federated Learning via Cluster-Cycling
by: Chen, Cheng, et al.
Published: (2020)
by: Chen, Cheng, et al.
Published: (2020)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
by: Bartoldson, Brian, et al.
Published: (2025)
by: Bartoldson, Brian, et al.
Published: (2025)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
End-to-End Mesh Optimization of a Hybrid Deep Learning Black-Box PDE Solver
by: Ma, Shaocong, et al.
Published: (2024)
by: Ma, Shaocong, et al.
Published: (2024)
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
by: Zheng, Haizhong, et al.
Published: (2026)
by: Zheng, Haizhong, et al.
Published: (2026)
Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis
by: Yang, Hongru, et al.
Published: (2024)
by: Yang, Hongru, et al.
Published: (2024)
Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
by: Cheng, Sitao, et al.
Published: (2025)
by: Cheng, Sitao, et al.
Published: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
Improving Robustness In Sparse Autoencoders via Masked Regularization
by: Narayanaswamy, Vivek, et al.
Published: (2026)
by: Narayanaswamy, Vivek, et al.
Published: (2026)
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
EchoRL: Reinforcement Learning via Rollout Echoing
by: Bi, Jinhe, et al.
Published: (2026)
by: Bi, Jinhe, et al.
Published: (2026)
Learning When to Act: Communication-Efficient Reinforcement Learning via Run-Time Assurance
by: Haroon, Adam, et al.
Published: (2026)
by: Haroon, Adam, et al.
Published: (2026)
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
by: Pal, Soumyadeep, et al.
Published: (2025)
by: Pal, Soumyadeep, et al.
Published: (2025)
Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
Debate Only When Necessary: Adaptive Multiagent Collaboration for Efficient LLM Reasoning
by: Eo, Sugyeong, et al.
Published: (2025)
by: Eo, Sugyeong, et al.
Published: (2025)
RoRecomp: Enhancing Reasoning Efficiency via Rollout Response Recomposition in Reinforcement Learning
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
ActSafe: Active Exploration with Safety Constraints for Reinforcement Learning
by: As, Yarden, et al.
Published: (2024)
by: As, Yarden, et al.
Published: (2024)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
by: Li, Yuhang, et al.
Published: (2026)
by: Li, Yuhang, et al.
Published: (2026)
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
by: Zhao, Jiawei, et al.
Published: (2024)
by: Zhao, Jiawei, et al.
Published: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
by: McLeish, Sean, et al.
Published: (2025)
by: McLeish, Sean, et al.
Published: (2025)
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
by: Onyame, Eric, et al.
Published: (2026)
by: Onyame, Eric, et al.
Published: (2026)
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
by: Yan, Kaizhuo, et al.
Published: (2025)
by: Yan, Kaizhuo, et al.
Published: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
by: McLeish, Sean, et al.
Published: (2024)
by: McLeish, Sean, et al.
Published: (2024)
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
by: Hong, Junyuan, et al.
Published: (2024)
by: Hong, Junyuan, et al.
Published: (2024)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Similar Items
-
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
by: Zheng, Haizhong, et al.
Published: (2024) -
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
by: Bartoldson, Brian R., et al.
Published: (2024) -
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
by: Wang, Zeyu, et al.
Published: (2025) -
Prosperity before Collapse: How Far Can Off-Policy RL Reach with Stale Data on LLMs?
by: Zheng, Haizhong, et al.
Published: (2025) -
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
by: McDonald, Tavish, et al.
Published: (2025)