Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Shenghua, Xia, Tian, Zhou, Xuan, Wei, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ToolRL: Reward is All Tool Learning Needs
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
von: Qian, Cheng, et al.
Veröffentlicht: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025)
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026)
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024)
von: Li, Junyou, et al.
Veröffentlicht: (2024)
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
Tensor Product Attention Is All You Need
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
von: Seshadri, Amrit Diggavi
Veröffentlicht: (2025)
Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable Rewards
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
von: Nguyen, Hieu Trung, et al.
Veröffentlicht: (2026)
Synthetic Data RL: Task Definition Is All You Need
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
von: Guo, Yiduo, et al.
Veröffentlicht: (2025)
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning
von: Li, Junsong, et al.
Veröffentlicht: (2025)
von: Li, Junsong, et al.
Veröffentlicht: (2025)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
von: Gan, Chunjing, et al.
Veröffentlicht: (2024)
von: Gan, Chunjing, et al.
Veröffentlicht: (2024)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
von: Song, Kefan, et al.
Veröffentlicht: (2025)
von: Song, Kefan, et al.
Veröffentlicht: (2025)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
von: Steinmetz, Cody, et al.
Veröffentlicht: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
von: Gomaa, Eyad, et al.
Veröffentlicht: (2024)
von: Gomaa, Eyad, et al.
Veröffentlicht: (2024)
All You Need is One: Capsule Prompt Tuning with a Single Vector
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
Easy Samples Are All You Need: Self-Evolving LLMs via Data-Efficient Reinforcement Learning
von: Yu, Zhiyin, et al.
Veröffentlicht: (2026)
von: Yu, Zhiyin, et al.
Veröffentlicht: (2026)
SecEncoder: Logs are All You Need in Security
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
von: Bulut, Muhammed Fatih, et al.
Veröffentlicht: (2024)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2025)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
von: Lu, Zhicong, et al.
Veröffentlicht: (2026)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
von: Sun, Lin, et al.
Veröffentlicht: (2025)
von: Sun, Lin, et al.
Veröffentlicht: (2025)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2024)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2024)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
What Matters in Transformers? Not All Attention is Needed
von: He, Shwai, et al.
Veröffentlicht: (2024)
von: He, Shwai, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Conditional Expectation Reward
von: Xiao, Changyi, et al.
Veröffentlicht: (2026)
von: Xiao, Changyi, et al.
Veröffentlicht: (2026)
The Lessons of Developing Process Reward Models in Mathematical Reasoning
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenru, et al.
Veröffentlicht: (2025)
Is Exploration All You Need? Effective Exploration Characteristics for Transfer in Reinforcement Learning
von: Balloch, Jonathan C., et al.
Veröffentlicht: (2024)
von: Balloch, Jonathan C., et al.
Veröffentlicht: (2024)
A Note on Hybrid Online Reinforcement and Imitation Learning for LLMs: Formulations and Algorithms
von: Li, Yingru, et al.
Veröffentlicht: (2025)
von: Li, Yingru, et al.
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs
von: Xie, Yutao, et al.
Veröffentlicht: (2026)
von: Xie, Yutao, et al.
Veröffentlicht: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
von: Ye, Zhiling, et al.
Veröffentlicht: (2025)
Evaluating Robustness of Reward Models for Mathematical Reasoning
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2024)
Attention is All You Need Until You Need Retention
von: Yaslioglu, M. Murat
Veröffentlicht: (2025)
von: Yaslioglu, M. Murat
Veröffentlicht: (2025)
T-REG: Preference Optimization with Token-Level Reward Regularization
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2024)
More Compute Is What You Need
von: Guo, Zhen
Veröffentlicht: (2024)
von: Guo, Zhen
Veröffentlicht: (2024)
Process Reinforcement through Implicit Rewards
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
Decoupling Reasoning and Confidence: Resurrecting Calibration in Reinforcement Learning from Verifiable Rewards
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
von: Ma, Zhengzhao, et al.
Veröffentlicht: (2026)
Hedging Is Not All You Need: A Simple Baseline for Online Learning Under Haphazard Inputs
von: Buckchash, Himanshu, et al.
Veröffentlicht: (2024)
von: Buckchash, Himanshu, et al.
Veröffentlicht: (2024)
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026)
von: Xia, Yu, et al.
Veröffentlicht: (2026)
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
von: Park, Jungsoo, et al.
Veröffentlicht: (2026)
von: Park, Jungsoo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ToolRL: Reward is All Tool Learning Needs
von: Qian, Cheng, et al.
Veröffentlicht: (2025) -
Attention Is All You Need for KV Cache in Diffusion LLMs
von: Nguyen-Tri, Quan, et al.
Veröffentlicht: (2025) -
Attention Smoothing Is All You Need For Unlearning
von: Zade, Saleh Zare, et al.
Veröffentlicht: (2026) -
More Agents Is All You Need
von: Li, Junyou, et al.
Veröffentlicht: (2024) -
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)