Learning to Reason without External Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Xuandong, Kang, Zhewei, Feng, Aosong, Levine, Sergey, Song, Dawn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Natural Language Actor-Critic: Scalable Off-Policy Learning in Language Space
von: Hong, Joey, et al.
Veröffentlicht: (2025)
von: Hong, Joey, et al.
Veröffentlicht: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
von: Lyu, Chengqi, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Semantic Rewards Enables Low-Resource Language Expansion without Alignment Tax
von: Su, Zeli, et al.
Veröffentlicht: (2026)
von: Su, Zeli, et al.
Veröffentlicht: (2026)
Learning to Assist Humans without Inferring Rewards
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
von: Myers, Vivek, et al.
Veröffentlicht: (2024)
Free Process Rewards without Process Labels
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
von: Bao, Keqin, et al.
Veröffentlicht: (2025)
von: Bao, Keqin, et al.
Veröffentlicht: (2025)
Making Bias Non-Predictive: Training Robust LLM Reasoning via Reinforcement Learning
von: Wang, Qian, et al.
Veröffentlicht: (2026)
von: Wang, Qian, et al.
Veröffentlicht: (2026)
$π$-Play: Multi-Agent Self-Play via Privileged Self-Distillation without External Data
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yaocheng, et al.
Veröffentlicht: (2026)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Scaling Reasoning without Attention
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
von: Zhao, Xueliang, et al.
Veröffentlicht: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Scaling Test-Time Compute Without Verification or RL is Suboptimal
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
von: Setlur, Amrith, et al.
Veröffentlicht: (2025)
Predicting Emergent Capabilities by Finetuning
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
von: Snell, Charlie, et al.
Veröffentlicht: (2024)
Reward-SQL: Boosting Text-to-SQL via Stepwise Reasoning and Process-Supervised Rewards
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
von: Zhang, Yuxin, et al.
Veröffentlicht: (2025)
GRAM-R$^2$: Self-Training Generative Foundation Reward Models for Reward Reasoning
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
von: Kang, Katie, et al.
Veröffentlicht: (2024)
von: Kang, Katie, et al.
Veröffentlicht: (2024)
Self-Sovereign Agent
von: Qu, Wenjie, et al.
Veröffentlicht: (2026)
von: Qu, Wenjie, et al.
Veröffentlicht: (2026)
Reinforcing General Reasoning without Verifiers
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
von: Crispino, Nicholas, et al.
Veröffentlicht: (2023)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
von: Liu, Wei, et al.
Veröffentlicht: (2025)
von: Liu, Wei, et al.
Veröffentlicht: (2025)
REASONING GYM: Reasoning Environments for Reinforcement Learning with Verifiable Rewards
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
von: Stojanovski, Zafir, et al.
Veröffentlicht: (2025)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
von: Setlur, Amrith, et al.
Veröffentlicht: (2024)
ConfClip: Confidence-Weighted and Clipped Reward for Reinforcement Learning in LLMs
von: Zhang, Bonan, et al.
Veröffentlicht: (2025)
von: Zhang, Bonan, et al.
Veröffentlicht: (2025)
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
von: Wu, Mian, et al.
Veröffentlicht: (2025)
von: Wu, Mian, et al.
Veröffentlicht: (2025)
ATLAS: Adaptive Test-Time Latent Steering with External Verifiers for Enhancing LLMs Reasoning
von: Nguyen, Tuc, et al.
Veröffentlicht: (2026)
von: Nguyen, Tuc, et al.
Veröffentlicht: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
von: Tang, Xinyu, et al.
Veröffentlicht: (2025)
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation
von: Min, Do June, et al.
Veröffentlicht: (2024)
von: Min, Do June, et al.
Veröffentlicht: (2024)
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026)
von: Xia, Yu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025) -
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025) -
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025) -
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025) -
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)