Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Kongcheng, Yao, Qi, Liu, Shunyu, Wang, Yingjie, Lai, Baisheng, Ye, Jieping, Song, Mingli, Tao, Dacheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Reasoning with Reinforced Functional Token Tuning
par: Zhang, Kongcheng, et autres
Publié: (2025)
par: Zhang, Kongcheng, et autres
Publié: (2025)
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
par: Fang, Wenkai, et autres
Publié: (2025)
par: Fang, Wenkai, et autres
Publié: (2025)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
par: Zhang, Kongcheng, et autres
Publié: (2025)
par: Zhang, Kongcheng, et autres
Publié: (2025)
Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs
par: Zhang, Wenjian, et autres
Publié: (2026)
par: Zhang, Wenjian, et autres
Publié: (2026)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
par: Wan, Guangya, et autres
Publié: (2024)
par: Wan, Guangya, et autres
Publié: (2024)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
par: Zhang, Jingyi, et autres
Publié: (2025)
par: Zhang, Jingyi, et autres
Publié: (2025)
Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
par: Zhou, Yang, et autres
Publié: (2025)
par: Zhou, Yang, et autres
Publié: (2025)
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
par: Zhang, Junjie, et autres
Publié: (2025)
par: Zhang, Junjie, et autres
Publié: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
par: Ye, Zhiling, et autres
Publié: (2025)
par: Ye, Zhiling, et autres
Publié: (2025)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
par: Zhao, Qingfei, et autres
Publié: (2025)
par: Zhao, Qingfei, et autres
Publié: (2025)
Towards Efficient LLM-aware Heterogeneous Graph Learning
par: Li, Wenda, et autres
Publié: (2025)
par: Li, Wenda, et autres
Publié: (2025)
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning
par: Ahmadi, Arash, et autres
Publié: (2026)
par: Ahmadi, Arash, et autres
Publié: (2026)
Intra-Trajectory Consistency for Reward Modeling
par: Zhou, Chaoyang, et autres
Publié: (2025)
par: Zhou, Chaoyang, et autres
Publié: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
par: Wang, Yibo, et autres
Publié: (2026)
par: Wang, Yibo, et autres
Publié: (2026)
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
par: Liu, Kaiyuan, et autres
Publié: (2025)
par: Liu, Kaiyuan, et autres
Publié: (2025)
Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing
par: Wu, Yuchen, et autres
Publié: (2025)
par: Wu, Yuchen, et autres
Publié: (2025)
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
par: Chen, Yanyu, et autres
Publié: (2026)
par: Chen, Yanyu, et autres
Publié: (2026)
Self-Consistency Boosts Calibration for Math Reasoning
par: Wang, Ante, et autres
Publié: (2024)
par: Wang, Ante, et autres
Publié: (2024)
Better, Faster: Harnessing Self-Improvement in Large Reasoning Models
par: Zhong, Qihuang, et autres
Publié: (2026)
par: Zhong, Qihuang, et autres
Publié: (2026)
CREAM: Consistency Regularized Self-Rewarding Language Models
par: Wang, Zhaoyang, et autres
Publié: (2024)
par: Wang, Zhaoyang, et autres
Publié: (2024)
Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
par: Liu, Kaiyuan, et autres
Publié: (2025)
par: Liu, Kaiyuan, et autres
Publié: (2025)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
par: Liu, Qihao, et autres
Publié: (2025)
par: Liu, Qihao, et autres
Publié: (2025)
Learning to Reason Across Parallel Samples for LLM Reasoning
par: Qi, Jianing, et autres
Publié: (2025)
par: Qi, Jianing, et autres
Publié: (2025)
Bridging Internal Probability and Self-Consistency for Effective and Efficient LLM Reasoning
par: Zhou, Zhi, et autres
Publié: (2025)
par: Zhou, Zhi, et autres
Publié: (2025)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
par: Wen, Xumeng, et autres
Publié: (2025)
par: Wen, Xumeng, et autres
Publié: (2025)
RLEP: Reinforcement Learning with Experience Replay for LLM Reasoning
par: Zhang, Hongzhi, et autres
Publié: (2025)
par: Zhang, Hongzhi, et autres
Publié: (2025)
Self-Reflective Planning with Knowledge Graphs: Enhancing LLM Reasoning Reliability for Question Answering
par: Zhu, Jiajun, et autres
Publié: (2025)
par: Zhu, Jiajun, et autres
Publié: (2025)
Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood Consistency
par: Xu, Haoming, et autres
Publié: (2026)
par: Xu, Haoming, et autres
Publié: (2026)
FreePRM: Training Process Reward Models Without Ground Truth Process Labels
par: Sun, Lin, et autres
Publié: (2025)
par: Sun, Lin, et autres
Publié: (2025)
WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
par: Qi, Zehan, et autres
Publié: (2024)
par: Qi, Zehan, et autres
Publié: (2024)
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning
par: Song, Jiwon, et autres
Publié: (2025)
par: Song, Jiwon, et autres
Publié: (2025)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
par: Wu, Tong, et autres
Publié: (2025)
par: Wu, Tong, et autres
Publié: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
par: Chen, Sirui, et autres
Publié: (2026)
par: Chen, Sirui, et autres
Publié: (2026)
Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
par: Xiong, Juming, et autres
Publié: (2026)
par: Xiong, Juming, et autres
Publié: (2026)
Detoxifying Large Language Models via Autoregressive Reward Guided Representation Editing
par: Xiao, Yisong, et autres
Publié: (2025)
par: Xiao, Yisong, et autres
Publié: (2025)
SSR-Zero: Simple Self-Rewarding Reinforcement Learning for Machine Translation
par: Yang, Wenjie, et autres
Publié: (2025)
par: Yang, Wenjie, et autres
Publié: (2025)
Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models
par: Nie, Shuo, et autres
Publié: (2026)
par: Nie, Shuo, et autres
Publié: (2026)
TriAlign: Towards Universal Truth Consistency in Personalized LLM Alignment
par: Nguyen, Thi-Nhung, et autres
Publié: (2026)
par: Nguyen, Thi-Nhung, et autres
Publié: (2026)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
par: Song, Mingyang, et autres
Publié: (2025)
par: Song, Mingyang, et autres
Publié: (2025)
On the Step Length Confounding in LLM Reasoning Data Selection
par: Wang, Bing, et autres
Publié: (2026)
par: Wang, Bing, et autres
Publié: (2026)
Documents similaires
-
Reasoning with Reinforced Functional Token Tuning
par: Zhang, Kongcheng, et autres
Publié: (2025) -
SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data
par: Fang, Wenkai, et autres
Publié: (2025) -
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
par: Zhang, Kongcheng, et autres
Publié: (2025) -
Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs
par: Zhang, Wenjian, et autres
Publié: (2026) -
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
par: Wan, Guangya, et autres
Publié: (2024)