Self-rewarding correction for mathematical reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiong, Wei, Zhang, Hanning, Ye, Chenlu, Chen, Lichang, Jiang, Nan, Zhang, Tong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
von: Xiong, Wei, et al.
Veröffentlicht: (2023)
von: Xiong, Wei, et al.
Veröffentlicht: (2023)
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026)
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
von: Yao, Jiarui, et al.
Veröffentlicht: (2025)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Streaming Looking Ahead with Token-level Self-reward
von: Zhang, Hongming, et al.
Veröffentlicht: (2025)
von: Zhang, Hongming, et al.
Veröffentlicht: (2025)
Math Takes Two: A test for emergent mathematical reasoning in communication
von: Cooper, Michael, et al.
Veröffentlicht: (2026)
von: Cooper, Michael, et al.
Veröffentlicht: (2026)
Towards Better Generalization via Distributional Input Projection Network
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Adaptive Layerwise Perturbation: Unifying Off-Policy Corrections for LLM RL
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
von: Ye, Chenlu, et al.
Veröffentlicht: (2026)
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Episodic Reinforcement Learning with Expanded State-reward Space
von: Liang, Dayang, et al.
Veröffentlicht: (2024)
von: Liang, Dayang, et al.
Veröffentlicht: (2024)
Reinforcement Learning in hyperbolic space for multi-step reasoning
von: Xu, Tao, et al.
Veröffentlicht: (2025)
von: Xu, Tao, et al.
Veröffentlicht: (2025)
LLMs cannot find reasoning errors, but can correct them given the error location
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
von: Tyen, Gladys, et al.
Veröffentlicht: (2023)
Noise-based reward-modulated learning
von: Fernández, Jesús García, et al.
Veröffentlicht: (2025)
von: Fernández, Jesús García, et al.
Veröffentlicht: (2025)
Active teacher selection for reward learning
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
von: Freedman, Rachel, et al.
Veröffentlicht: (2023)
Online Iterative Reinforcement Learning from Human Feedback with General Preference Model
von: Ye, Chenlu, et al.
Veröffentlicht: (2024)
von: Ye, Chenlu, et al.
Veröffentlicht: (2024)
ProvMind: Provenance-grounded reasoning for materials synthesis
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
The impact of intrinsic rewards on exploration in Reinforcement Learning
von: Kayal, Aya, et al.
Veröffentlicht: (2025)
von: Kayal, Aya, et al.
Veröffentlicht: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
Understanding Overadaptation in Supervised Fine-Tuning: The Role of Ensemble Methods
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
von: Hao, Yifan, et al.
Veröffentlicht: (2025)
Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2025)
On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
von: Zhang, Yuheng, et al.
Veröffentlicht: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
EVAL: EigenVector-based Average-reward Learning
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
von: Adamczyk, Jacob, et al.
Veröffentlicht: (2025)
Exploring System 1 and 2 communication for latent reasoning in LLMs
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2025)
von: Coda-Forno, Julian, et al.
Veröffentlicht: (2025)
Risk-averse Total-reward MDPs with ERM and EVaR
von: Su, Xihong, et al.
Veröffentlicht: (2024)
von: Su, Xihong, et al.
Veröffentlicht: (2024)
reward-lens: A Mechanistic Interpretability Library for Reward Models
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
von: Nadaf, Mohammed Suhail B
Veröffentlicht: (2026)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision Processes
von: Ye, Chenlu, et al.
Veröffentlicht: (2022)
von: Ye, Chenlu, et al.
Veröffentlicht: (2022)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
Divergence-Augmented Policy Optimization
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
A Minimalist Approach to LLM Reasoning: from Rejection Sampling to Reinforce
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
Selecting Belief-State Approximations in Simulators with Latent States
von: Jiang, Nan
Veröffentlicht: (2025)
von: Jiang, Nan
Veröffentlicht: (2025)
A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning
von: Jiang, Nan
Veröffentlicht: (2024)
von: Jiang, Nan
Veröffentlicht: (2024)
Art and Science of Quantizing Large-Scale Models: A Comprehensive Overview
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
von: Wang, Yanshu, et al.
Veröffentlicht: (2024)
Zero-Incentive Dynamics: a look at reward sparsity through the lens of unrewarded subgoals
von: Molinghen, Yannick, et al.
Veröffentlicht: (2025)
von: Molinghen, Yannick, et al.
Veröffentlicht: (2025)
MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
von: Ding, Chenlu, et al.
Veröffentlicht: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
von: Shen, Junhong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
von: Xiong, Wei, et al.
Veröffentlicht: (2023) -
Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective
von: Zhang, Yuheng, et al.
Veröffentlicht: (2026) -
Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL
von: Yao, Jiarui, et al.
Veröffentlicht: (2025) -
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
von: Xiong, Wei, et al.
Veröffentlicht: (2025) -
Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models
von: Hao, Yifan, et al.
Veröffentlicht: (2025)