$\textbf{Re}^{2}$: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Pinzheng, Xu, Shuli, Li, Juntao, Luo, Yu, Li, Dong, Hao, Jianye, Zhang, Min |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
MODULI: Unlocking Preference Generalization via Diffusion Models for Offline Multi-Objective Reinforcement Learning
by: Yuan, Yifu, et al.
Published: (2024)
by: Yuan, Yifu, et al.
Published: (2024)
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
by: Chen, Mingyang, et al.
Published: (2025)
by: Chen, Mingyang, et al.
Published: (2025)
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025)
by: Wang, Pinzheng, et al.
Published: (2025)
ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning
by: Huang, Jiani, et al.
Published: (2026)
by: Huang, Jiani, et al.
Published: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
by: Xu, Wanghan, et al.
Published: (2026)
by: Xu, Wanghan, et al.
Published: (2026)
Unlocking Recursive Thinking of LLMs: Alignment via Refinement
by: Zhang, Haoke, et al.
Published: (2025)
by: Zhang, Haoke, et al.
Published: (2025)
ReCreate: Reasoning and Creating Domain Agents Driven by Experience
by: Hao, Zhezheng, et al.
Published: (2026)
by: Hao, Zhezheng, et al.
Published: (2026)
Ratio-Variance Regularized Policy Optimization for Efficient LLM Fine-tuning
by: Luo, Yu, et al.
Published: (2026)
by: Luo, Yu, et al.
Published: (2026)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
by: Xu, Mingjun, et al.
Published: (2025)
by: Xu, Mingjun, et al.
Published: (2025)
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
by: Wang, Hong, et al.
Published: (2025)
by: Wang, Hong, et al.
Published: (2025)
Re2LLM: Reflective Reinforcement Large Language Model for Session-based Recommendation
by: Wang, Ziyan, et al.
Published: (2024)
by: Wang, Ziyan, et al.
Published: (2024)
ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Beyond ReAct: A Planner-Centric Framework for Complex Tool-Augmented LLM Reasoning
by: Wei, Xiaolong, et al.
Published: (2025)
by: Wei, Xiaolong, et al.
Published: (2025)
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning
by: Hao, Qianyue, et al.
Published: (2025)
by: Hao, Qianyue, et al.
Published: (2025)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
by: Zhang, Zhaoyang, et al.
Published: (2026)
by: Zhang, Zhaoyang, et al.
Published: (2026)
Generative Reasoning Re-ranker
by: Liang, Mingfu, et al.
Published: (2026)
by: Liang, Mingfu, et al.
Published: (2026)
ReSS: Learning Reasoning Models for Tabular Data Prediction via Symbolic Scaffold
by: Yi, Chenlang, et al.
Published: (2026)
by: Yi, Chenlang, et al.
Published: (2026)
Unleashing LLM Reasoning Capability via Scalable Question Synthesis from Scratch
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
by: Fan, Lishui, et al.
Published: (2025)
by: Fan, Lishui, et al.
Published: (2025)
OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents
by: Yu, Sheldon, et al.
Published: (2026)
by: Yu, Sheldon, et al.
Published: (2026)
Focused ReAct: Improving ReAct through Reiterate and Early Stop
by: Li, Shuoqiu, et al.
Published: (2024)
by: Li, Shuoqiu, et al.
Published: (2024)
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent
by: Wang, Prince Zizhuang, et al.
Published: (2026)
by: Wang, Prince Zizhuang, et al.
Published: (2026)
Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning
by: Wu, Qingyuan, et al.
Published: (2025)
by: Wu, Qingyuan, et al.
Published: (2025)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
by: Wang, Huaijie, et al.
Published: (2024)
by: Wang, Huaijie, et al.
Published: (2024)
HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking
by: Gui, Runquan, et al.
Published: (2025)
by: Gui, Runquan, et al.
Published: (2025)
ReLA: Representation Learning and Aggregation for Job Scheduling with Reinforcement Learning
by: Kwan, Zhengyi, et al.
Published: (2026)
by: Kwan, Zhengyi, et al.
Published: (2026)
Jailbreak Instruction-Tuned LLMs via end-of-sentence MLP Re-weighting
by: Luo, Yifan, et al.
Published: (2024)
by: Luo, Yifan, et al.
Published: (2024)
MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction
by: Tang, Zitian, et al.
Published: (2026)
by: Tang, Zitian, et al.
Published: (2026)
On Information Self-Locking in Reinforcement Learning for Active Reasoning of LLM agents
by: Zou, Deyu, et al.
Published: (2026)
by: Zou, Deyu, et al.
Published: (2026)
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
by: Dai, Gaole, et al.
Published: (2025)
by: Dai, Gaole, et al.
Published: (2025)
Self-ReSET: Learning to Self-Recover from Unsafe Reasoning Trajectories
by: Zhang, Dongcheng, et al.
Published: (2026)
by: Zhang, Dongcheng, et al.
Published: (2026)
SATURN: SAT-based Reinforcement Learning to Unleash LLMs Reasoning
by: Liu, Huanyu, et al.
Published: (2025)
by: Liu, Huanyu, et al.
Published: (2025)
LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
by: Ning, Yansong, et al.
Published: (2025)
by: Ning, Yansong, et al.
Published: (2025)
CoRe-Code: Collaborative Reinforcement Learning for Code Generation
by: Dou, Zhihao, et al.
Published: (2026)
by: Dou, Zhihao, et al.
Published: (2026)
Formula-R1: Incentivizing LLM Reasoning over Complex Tables with Numerical Computation via Formula-Driven Reinforcement Learning
by: Cao, Lang, et al.
Published: (2025)
by: Cao, Lang, et al.
Published: (2025)
KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance
by: Yu, Linhao, et al.
Published: (2026)
by: Yu, Linhao, et al.
Published: (2026)
Similar Items
-
Improving Rationality in the Reasoning Process of Language Models through Self-playing Game
by: Wang, Pinzheng, et al.
Published: (2025) -
MODULI: Unlocking Preference Generalization via Diffusion Models for Offline Multi-Objective Reinforcement Learning
by: Yuan, Yifu, et al.
Published: (2024) -
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
by: Chen, Mingyang, et al.
Published: (2025) -
Revealing and Mitigating Over-Attention in Knowledge Editing
by: Wang, Pinzheng, et al.
Published: (2025) -
ReRec: Reasoning-Augmented LLM-based Recommendation Assistant via Reinforcement Fine-tuning
by: Huang, Jiani, et al.
Published: (2026)