Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Turpin, Miles, Arditi, Andy, Li, Marvin, Benton, Joe, Michael, Julian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
von: Chua, James, et al.
Veröffentlicht: (2024)
von: Chua, James, et al.
Veröffentlicht: (2024)
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
von: Wang, Jiecong, et al.
Veröffentlicht: (2026)
von: Wang, Jiecong, et al.
Veröffentlicht: (2026)
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025)
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
von: Xiong, Xuan, et al.
Veröffentlicht: (2026)
von: Xiong, Xuan, et al.
Veröffentlicht: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2025)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
Explainable Chain-of-Thought Reasoning: An Empirical Analysis on State-Aware Reasoning Dynamics
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
von: Yu, Sheldon, et al.
Veröffentlicht: (2025)
Spontaneous Reward Hacking in Iterative Self-Refinement
von: Pan, Jane, et al.
Veröffentlicht: (2024)
von: Pan, Jane, et al.
Veröffentlicht: (2024)
CSRP: Chain-of-Thought Reasoning for Chinese Text Correction via Reinforcement Learning with Efficiency-Aware Rewards
von: Tian, Wei, et al.
Veröffentlicht: (2026)
von: Tian, Wei, et al.
Veröffentlicht: (2026)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
Latent Chain-of-Thought for Visual Reasoning
von: Sun, Guohao, et al.
Veröffentlicht: (2025)
von: Sun, Guohao, et al.
Veröffentlicht: (2025)
The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
von: Yang, Xinglong, et al.
Veröffentlicht: (2025)
Fractured Chain-of-Thought Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
How does Chain of Thought Think? Mechanistic Interpretability of Chain-of-Thought Reasoning with Sparse Autoencoding
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning in Language Models
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
Markov Chain of Thought for Efficient Mathematical Reasoning
von: Yang, Wen, et al.
Veröffentlicht: (2024)
von: Yang, Wen, et al.
Veröffentlicht: (2024)
GRACE: Discriminator-Guided Chain-of-Thought Reasoning
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2023)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2023)
LaRS: Latent Reasoning Skills for Chain-of-Thought Reasoning
von: Xu, Zifan, et al.
Veröffentlicht: (2023)
von: Xu, Zifan, et al.
Veröffentlicht: (2023)
Inverse Scaling in Test-Time Compute
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2025)
von: Gema, Aryo Pradipta, et al.
Veröffentlicht: (2025)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
von: Zhu, Zihao, et al.
Veröffentlicht: (2025)
LLMs can Find Mathematical Reasoning Mistakes by Pedagogical Chain-of-Thought
von: Jiang, Zhuoxuan, et al.
Veröffentlicht: (2024)
von: Jiang, Zhuoxuan, et al.
Veröffentlicht: (2024)
Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs
von: Xia, Yu, et al.
Veröffentlicht: (2024)
von: Xia, Yu, et al.
Veröffentlicht: (2024)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering
von: Wu, Yexin, et al.
Veröffentlicht: (2024)
von: Wu, Yexin, et al.
Veröffentlicht: (2024)
To Reason or Not to: Selective Chain-of-Thought in Medical Question Answering
von: Zhan, Zaifu, et al.
Veröffentlicht: (2026)
von: Zhan, Zaifu, et al.
Veröffentlicht: (2026)
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
von: Chen, Qiguang, et al.
Veröffentlicht: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
Fine-Tuned Thoughts: Leveraging Chain-of-Thought Reasoning for Industrial Asset Health Monitoring
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
von: Lin, Shuxin, et al.
Veröffentlicht: (2025)
Monitoring Emergent Reward Hacking During Generation via Internal Activations
von: Wilhelm, Patrick, et al.
Veröffentlicht: (2026)
von: Wilhelm, Patrick, et al.
Veröffentlicht: (2026)
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
von: Yuan, Aojie, et al.
Veröffentlicht: (2026)
CoAT: Chain-of-Associated-Thoughts Framework for Enhancing Large Language Models Reasoning
von: Pan, Jianfeng, et al.
Veröffentlicht: (2025)
von: Pan, Jianfeng, et al.
Veröffentlicht: (2025)
Expanding Reasoning Potential in Foundation Model by Learning Diverse Chains of Thought Patterns
von: Zhang, Xuemiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xuemiao, et al.
Veröffentlicht: (2025)
Injecting Salesperson's Dialogue Strategies in Large Language Models with Chain-of-Thought Reasoning
von: Chang, Wen-Yu, et al.
Veröffentlicht: (2024)
von: Chang, Wen-Yu, et al.
Veröffentlicht: (2024)
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
von: Yin, Maxwell J., et al.
Veröffentlicht: (2025)
von: Yin, Maxwell J., et al.
Veröffentlicht: (2025)
ConMax: Confidence-Maximizing Compression for Efficient Chain-of-Thought Reasoning
von: Hu, Minda, et al.
Veröffentlicht: (2026)
von: Hu, Minda, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
von: Chua, James, et al.
Veröffentlicht: (2024) -
Latent Chain-of-Thought as Planning: Decoupling Reasoning from Verbalization
von: Wang, Jiecong, et al.
Veröffentlicht: (2026) -
Where Do Reasoning Models Refuse?
von: Yamaguchi, Kureha, et al.
Veröffentlicht: (2025) -
ETR: Entropy Trend Reward for Efficient Chain-of-Thought Reasoning
von: Xiong, Xuan, et al.
Veröffentlicht: (2026) -
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)