Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertainty
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Chao, Wang, Yao, Liu, Mengqiao, Liang, Di, Han, Xingsheng, Liu, Peiyang, Wu, Xianjie, Lu, Chenyao, Jiang, Lei, Lu, Yu, Shi, Haibo, Liang, Shuang, Peng, Minlong, Salim, Flora D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
von: Xue, Chao, et al.
Veröffentlicht: (2026)
von: Xue, Chao, et al.
Veröffentlicht: (2026)
Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
von: Lin, Zekai, et al.
Veröffentlicht: (2026)
von: Lin, Zekai, et al.
Veröffentlicht: (2026)
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025)
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
von: Wang, Yao, et al.
Veröffentlicht: (2025)
von: Wang, Yao, et al.
Veröffentlicht: (2025)
Think Only When You Need with Large Hybrid-Reasoning Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
Mechanistic Indicators of Steering Effectiveness in Large Language Models
von: Jafari, Mehdi, et al.
Veröffentlicht: (2026)
von: Jafari, Mehdi, et al.
Veröffentlicht: (2026)
When You Need to Know
von: Brown, Flora Morris
Veröffentlicht: (1976)
von: Brown, Flora Morris
Veröffentlicht: (1976)
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models
von: Hong, Ilgee, et al.
Veröffentlicht: (2025)
von: Hong, Ilgee, et al.
Veröffentlicht: (2025)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Reinforcement Learning Enhanced Multi-hop Reasoning for Temporal Knowledge Question Answering
von: Wen, Wuzhenghong, et al.
Veröffentlicht: (2026)
von: Wen, Wuzhenghong, et al.
Veröffentlicht: (2026)
Whether Uncertainty Theory Can Enhance GDP Forecasting From Energy: A New Uncertain MIDAS Model
von: Yuxin Shi, et al.
Veröffentlicht: (2025)
von: Yuxin Shi, et al.
Veröffentlicht: (2025)
DeepLévy: Learning Heavy-Tailed Uncertainty in Highly Volatile Time Series
von: Yang, Yang, et al.
Veröffentlicht: (2026)
von: Yang, Yang, et al.
Veröffentlicht: (2026)
Uncertainty Quantification for Large Language Model Reward Learning under Heterogeneous Human Feedback
von: Liu, Pangpang, et al.
Veröffentlicht: (2025)
von: Liu, Pangpang, et al.
Veröffentlicht: (2025)
DPI: Exploiting Parameter Heterogeneity for Interference-Free Fine-Tuning
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2026)
Reward Model Routing in Alignment
von: Wu, Xinle, et al.
Veröffentlicht: (2025)
von: Wu, Xinle, et al.
Veröffentlicht: (2025)
MAPLE: Mobile App Prediction Leveraging Large Language Model Embeddings
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2023)
von: Khaokaew, Yonchanok, et al.
Veröffentlicht: (2023)
Who Stole Your Data? A Method for Detecting Unauthorized RAG Theft
von: Liu, Peiyang, et al.
Veröffentlicht: (2025)
von: Liu, Peiyang, et al.
Veröffentlicht: (2025)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
von: Liang, Yiming, et al.
Veröffentlicht: (2026)
Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents
von: Cai, Yuxuan, et al.
Veröffentlicht: (2026)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2026)
Take Only What You Need: Rank Minimization as an Implicit Forgetting Regularizer in Continual Learning
von: Lu, Haodong, et al.
Veröffentlicht: (2024)
von: Lu, Haodong, et al.
Veröffentlicht: (2024)
Large Language Models are Contrastive Reasoners
von: Yao, Liang
Veröffentlicht: (2024)
von: Yao, Liang
Veröffentlicht: (2024)
Optimal Uncertainty Quantification under General Moment Constraints on Input Subdomains
von: Jin, Rong, et al.
Veröffentlicht: (2025)
von: Jin, Rong, et al.
Veröffentlicht: (2025)
Double-Diffusion: ODE-Prior Accelerated Diffusion Models for Spatio-Temporal Graph Forecasting
von: Dong, Hanlin, et al.
Veröffentlicht: (2025)
von: Dong, Hanlin, et al.
Veröffentlicht: (2025)
A-UTE: Advection Informed, Uncertainty Aware Temperature Emulator
von: Saleem, Hira, et al.
Veröffentlicht: (2024)
von: Saleem, Hira, et al.
Veröffentlicht: (2024)
Large Language Models for Next Point-of-Interest Recommendation
von: Li, Peibo, et al.
Veröffentlicht: (2024)
von: Li, Peibo, et al.
Veröffentlicht: (2024)
When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
von: Chen, Baiyu, et al.
Veröffentlicht: (2025)
Decoder-Only LLMs are Better Controllers for Diffusion Models
von: Dong, Ziyi, et al.
Veröffentlicht: (2025)
von: Dong, Ziyi, et al.
Veröffentlicht: (2025)
STAR: Failure-Aware Markovian Routing for Multi-Agent Spatiotemporal Reasoning
von: Yang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Yang, Ruiyi, et al.
Veröffentlicht: (2026)
From XXLTraffic to EvoXXLTraffic: Scaling Traffic Forecasting to Sensor-Evolving Networks
von: Yin, Du, et al.
Veröffentlicht: (2026)
von: Yin, Du, et al.
Veröffentlicht: (2026)
XXLTraffic: Expanding and Extremely Long Traffic forecasting beyond test adaptation
von: Yin, Du, et al.
Veröffentlicht: (2024)
von: Yin, Du, et al.
Veröffentlicht: (2024)
HiT-JEPA: A Hierarchical Self-supervised Trajectory Embedding Framework for Similarity Computation
von: Li, Lihuan, et al.
Veröffentlicht: (2025)
von: Li, Lihuan, et al.
Veröffentlicht: (2025)
Large Language Model Reasoning Failures
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
von: Song, Peiyang, et al.
Veröffentlicht: (2026)
Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
Free(): Learning to Forget in Malloc-Only Reasoning Models
von: Zheng, Yilun, et al.
Veröffentlicht: (2026)
von: Zheng, Yilun, et al.
Veröffentlicht: (2026)
EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
von: Ao, Shuang, et al.
Veröffentlicht: (2025)
von: Ao, Shuang, et al.
Veröffentlicht: (2025)
ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents
von: Li, Zechen, et al.
Veröffentlicht: (2025)
von: Li, Zechen, et al.
Veröffentlicht: (2025)
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
von: Wang, Yuanfu, et al.
Veröffentlicht: (2026)
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
von: Yang, Yankai, et al.
Veröffentlicht: (2026)
von: Yang, Yankai, et al.
Veröffentlicht: (2026)
When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG
von: Li, Junchen, et al.
Veröffentlicht: (2026)
von: Li, Junchen, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Why Supervised Fine-Tuning Fails to Learn: A Systematic Study of Incomplete Learning in Large Language Models
von: Xue, Chao, et al.
Veröffentlicht: (2026) -
Parameter Importance is Not Static: Evolving Parameter Isolation for Supervised Fine-Tuning
von: Lin, Zekai, et al.
Veröffentlicht: (2026) -
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025) -
DeCoRL: Decoupling Reasoning Chains via Parallel Sub-Step Generation and Cascaded Reinforcement for Interpretable and Scalable RLHF
von: Gao, Ziyuan, et al.
Veröffentlicht: (2025) -
Not All Parameters Are Created Equal: Smart Isolation Boosts Fine-Tuning Performance
von: Wang, Yao, et al.
Veröffentlicht: (2025)