REA-RL: Reflection-Aware Online Reinforcement Learning for Efficient Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Deng, Hexuan, Jiao, Wenxiang, Liu, Xuebo, Rao, Jun, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
by: Deng, Hexuan, et al.
Published: (2024)
by: Deng, Hexuan, et al.
Published: (2024)
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
by: Rao, Jun, et al.
Published: (2025)
by: Rao, Jun, et al.
Published: (2025)
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
by: Yuan, Bo, et al.
Published: (2026)
by: Yuan, Bo, et al.
Published: (2026)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
by: Liu, Liangxin, et al.
Published: (2024)
by: Liu, Liangxin, et al.
Published: (2024)
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs
by: Zhou, Runlong, et al.
Published: (2024)
by: Zhou, Runlong, et al.
Published: (2024)
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models
by: Nie, Shuo, et al.
Published: (2026)
by: Nie, Shuo, et al.
Published: (2026)
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning
by: Shandilya, Shivam, et al.
Published: (2024)
by: Shandilya, Shivam, et al.
Published: (2024)
TemplateRL: Structured Template-Guided Reinforcement Learning for LLM Reasoning
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
by: Shi, Yucheng, et al.
Published: (2025)
by: Shi, Yucheng, et al.
Published: (2025)
UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
AQuilt: Weaving Logic and Self-Inspection into Low-Cost, High-Relevance Data Synthesis for Specialist LLMs
by: Ke, Xiaopeng, et al.
Published: (2025)
by: Ke, Xiaopeng, et al.
Published: (2025)
Beyond Markovian: Reflective Exploration via Bayes-Adaptive RL for LLM Reasoning
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
by: Peng, Keqin, et al.
Published: (2026)
by: Peng, Keqin, et al.
Published: (2026)
BroRL: Scaling Reinforcement Learning via Broadened Exploration
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
SPEC-RL: Accelerating On-Policy Reinforcement Learning with Speculative Rollouts
by: Liu, Bingshuai, et al.
Published: (2025)
by: Liu, Bingshuai, et al.
Published: (2025)
Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models
by: Liu, Runze, et al.
Published: (2025)
by: Liu, Runze, et al.
Published: (2025)
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
Flow-DPO: Improving LLM Mathematical Reasoning through Online Multi-Agent Learning
by: Deng, Yihe, et al.
Published: (2024)
by: Deng, Yihe, et al.
Published: (2024)
RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
by: Zha, Kaiwen, et al.
Published: (2025)
by: Zha, Kaiwen, et al.
Published: (2025)
Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
by: Kim, Junseok, et al.
Published: (2026)
by: Kim, Junseok, et al.
Published: (2026)
NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
by: Wang, Yinuo, et al.
Published: (2026)
by: Wang, Yinuo, et al.
Published: (2026)
Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learning
by: Liu, Hanbing, et al.
Published: (2025)
by: Liu, Hanbing, et al.
Published: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
TreeRL: LLM Reinforcement Learning with On-Policy Tree Search
by: Hou, Zhenyu, et al.
Published: (2025)
by: Hou, Zhenyu, et al.
Published: (2025)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
ConciseRL: Conciseness-Guided Reinforcement Learning for Efficient Reasoning Models
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2025)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
by: Qiu, Zhaopeng, et al.
Published: (2026)
by: Qiu, Zhaopeng, et al.
Published: (2026)
Balancing the Reasoning Load: Difficulty-Differentiated Policy Optimization with Length Redistribution for Efficient and Robust Reinforcement Learning
by: Xia, Yinan, et al.
Published: (2026)
by: Xia, Yinan, et al.
Published: (2026)
Meta-Reinforcement Learning with Self-Reflection for Agentic Search
by: Xiao, Teng, et al.
Published: (2026)
by: Xiao, Teng, et al.
Published: (2026)
Reinforce LLM Reasoning through Multi-Agent Reflection
by: Yuan, Yurun, et al.
Published: (2025)
by: Yuan, Yurun, et al.
Published: (2025)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
by: Yang, Rui, et al.
Published: (2026)
by: Yang, Rui, et al.
Published: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026)
by: Xu, Wanghan, et al.
Published: (2026)
Reinforcing General Reasoning without Verifiers
by: Zhou, Xiangxin, et al.
Published: (2025)
by: Zhou, Xiangxin, et al.
Published: (2025)
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
ExPO: Unlocking Hard Reasoning with Self-Explanation-Guided Reinforcement Learning
by: Zhou, Ruiyang, et al.
Published: (2025)
by: Zhou, Ruiyang, et al.
Published: (2025)
Stable and Efficient Single-Rollout RL for Multimodal Reasoning
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
by: Deng, Yihe, et al.
Published: (2025)
by: Deng, Yihe, et al.
Published: (2025)
Similar Items
-
NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
by: Deng, Hexuan, et al.
Published: (2024) -
DRPruning: Efficient Large Language Model Pruning through Distributionally Robust Optimization
by: Deng, Hexuan, et al.
Published: (2024) -
Dynamic Sampling that Adapts: Self-Aware Iterative Data Persistent Optimization for Mathematical Reasoning
by: Rao, Jun, et al.
Published: (2025) -
RouterKGQA: Specialized--General Model Routing for Constraint-Aware Knowledge Graph Question Answering
by: Yuan, Bo, et al.
Published: (2026) -
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
by: Liu, Liangxin, et al.
Published: (2024)