Likelihood-Based Reward Designs for General LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kwiatkowski, Ariel, Butt, Natasha, Labiad, Ismail, Kempe, Julia, Ollivier, Yann |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Soft Tokens, Hard Truths
von: Butt, Natasha, et al.
Veröffentlicht: (2025)
von: Butt, Natasha, et al.
Veröffentlicht: (2025)
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2026)
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2026)
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
von: Labiad, Ismail, et al.
Veröffentlicht: (2025)
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025)
von: Song, Yuda, et al.
Veröffentlicht: (2025)
PILAF: Optimal Human Preference Sampling for Reward Modeling
von: Feng, Yunzhen, et al.
Veröffentlicht: (2025)
von: Feng, Yunzhen, et al.
Veröffentlicht: (2025)
Unifying Tree Search Algorithm and Reward Design for LLM Reasoning: A Survey
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
On Designing Effective RL Reward at Training Time for LLM Reasoning
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2024)
AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
von: Jia, Mengzhao, et al.
Veröffentlicht: (2025)
von: Jia, Mengzhao, et al.
Veröffentlicht: (2025)
Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
von: Yuan, Youliang, et al.
Veröffentlicht: (2025)
Aligning Language Models Using Follow-up Likelihood as Reward Signal
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
GeneralThinker: Domain-General Reasoning through Likelihood-Guided Answer-Conditioned Optimization
von: Piao, Shengmin, et al.
Veröffentlicht: (2026)
von: Piao, Shengmin, et al.
Veröffentlicht: (2026)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
von: Bhardwaj, Dhrupad, et al.
Veröffentlicht: (2025)
von: Bhardwaj, Dhrupad, et al.
Veröffentlicht: (2025)
From Concepts to Components: Concept-Agnostic Attention Module Discovery in Transformers
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
von: Su, Jingtong, et al.
Veröffentlicht: (2025)
Mission Impossible: A Statistical Perspective on Jailbreaking LLMs
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
von: Su, Jingtong, et al.
Veröffentlicht: (2024)
Enhancing LLM Reasoning with Reward-guided Tree Search
von: Jiang, Jinhao, et al.
Veröffentlicht: (2024)
von: Jiang, Jinhao, et al.
Veröffentlicht: (2024)
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
PACR: Progressively Ascending Confidence Reward for LLM Reasoning
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
von: Yoon, Eunseop, et al.
Veröffentlicht: (2025)
CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models
von: Tan, Zhehao, et al.
Veröffentlicht: (2026)
von: Tan, Zhehao, et al.
Veröffentlicht: (2026)
Improving Data and Reward Design for Scientific Reasoning in Large Language Models
von: Chen, Zijie, et al.
Veröffentlicht: (2026)
von: Chen, Zijie, et al.
Veröffentlicht: (2026)
Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning
von: Ahmadi, Arash, et al.
Veröffentlicht: (2026)
von: Ahmadi, Arash, et al.
Veröffentlicht: (2026)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
von: Xu, Jundong, et al.
Veröffentlicht: (2025)
von: Xu, Jundong, et al.
Veröffentlicht: (2025)
Reward Reasoning Model
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
von: Guo, Jiaxin, et al.
Veröffentlicht: (2025)
A Survey on Progress in LLM Alignment from the Perspective of Reward Design
von: Ji, Miaomiao, et al.
Veröffentlicht: (2025)
von: Ji, Miaomiao, et al.
Veröffentlicht: (2025)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
von: Xu, Zhichao, et al.
Veröffentlicht: (2025)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
von: Zhao, Qingfei, et al.
Veröffentlicht: (2025)
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
von: Zhou, Hongli, et al.
Veröffentlicht: (2026)
Reward Design for Physical Reasoning in Vision-Language Models
von: Lilienthal, Derek, et al.
Veröffentlicht: (2026)
von: Lilienthal, Derek, et al.
Veröffentlicht: (2026)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Metacognition as Reward: Reinforcing LLM Reasoning via Knowledge and Regulation Signals
von: Chen, Sirui, et al.
Veröffentlicht: (2026)
von: Chen, Sirui, et al.
Veröffentlicht: (2026)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Reasoning Boosts Opinion Alignment in LLMs
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
ReasonGRM: Enhancing Generative Reward Models through Large Reasoning Models
von: Chen, Bin, et al.
Veröffentlicht: (2025)
von: Chen, Bin, et al.
Veröffentlicht: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood
von: Lin, Xingyu, et al.
Veröffentlicht: (2025)
von: Lin, Xingyu, et al.
Veröffentlicht: (2025)
From RAG to RICHES: Retrieval Interlaced with Sequence Generation
von: Jain, Palak, et al.
Veröffentlicht: (2024)
von: Jain, Palak, et al.
Veröffentlicht: (2024)
General-Reasoner: Advancing LLM Reasoning Across All Domains
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
von: Lan, Kai, et al.
Veröffentlicht: (2025)
von: Lan, Kai, et al.
Veröffentlicht: (2025)
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Sequence-Level Likelihood
von: Lin, Xingyu, et al.
Veröffentlicht: (2026)
von: Lin, Xingyu, et al.
Veröffentlicht: (2026)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
von: Zhang, Jianghangfan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Soft Tokens, Hard Truths
von: Butt, Natasha, et al.
Veröffentlicht: (2025) -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
von: Sundaram, Shobhita, et al.
Veröffentlicht: (2026) -
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
von: Labiad, Ismail, et al.
Veröffentlicht: (2025) -
Outcome-based Exploration for LLM Reasoning
von: Song, Yuda, et al.
Veröffentlicht: (2025) -
PILAF: Optimal Human Preference Sampling for Reward Modeling
von: Feng, Yunzhen, et al.
Veröffentlicht: (2025)