Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinpeng, Joshi, Nitish, Plank, Barbara, Angell, Rico, He, He |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
von: Si, Shengyun, et al.
Veröffentlicht: (2025)
Spontaneous Reward Hacking in Iterative Self-Refinement
von: Pan, Jane, et al.
Veröffentlicht: (2024)
von: Pan, Jane, et al.
Veröffentlicht: (2024)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024)
LLMs Are Prone to Fallacies in Causal Inference
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
von: Turpin, Miles, et al.
Veröffentlicht: (2025)
von: Turpin, Miles, et al.
Veröffentlicht: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Reward Shaping to Mitigate Reward Hacking in RLHF
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
von: Fu, Jiayi, et al.
Veröffentlicht: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
von: Mondorf, Philipp, et al.
Veröffentlicht: (2024)
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
von: He, Yanjie
Veröffentlicht: (2026)
von: He, Yanjie
Veröffentlicht: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2026)
LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Models
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
von: Rabern, Brian, et al.
Veröffentlicht: (2026)
Personas as a Way to Model Truthfulness in Language Models
von: Joshi, Nitish, et al.
Veröffentlicht: (2023)
von: Joshi, Nitish, et al.
Veröffentlicht: (2023)
Do LLMs Really Think Step-by-step In Implicit Reasoning?
von: Yu, Yijiong
Veröffentlicht: (2024)
von: Yu, Yijiong
Veröffentlicht: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
von: Chen, Lichang, et al.
Veröffentlicht: (2024)
Monitoring Emergent Reward Hacking During Generation via Internal Activations
von: Wilhelm, Patrick, et al.
Veröffentlicht: (2026)
von: Wilhelm, Patrick, et al.
Veröffentlicht: (2026)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
von: Wen, Xumeng, et al.
Veröffentlicht: (2025)
von: Wen, Xumeng, et al.
Veröffentlicht: (2025)
Thinking with DistilQwen: A Tale of Four Distilled Reasoning and Reward Model Series
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
von: Cai, Wenrui, et al.
Veröffentlicht: (2025)
Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2025)
von: Yueh-Han, Chen, et al.
Veröffentlicht: (2025)
Reasoning Models Can Be Effective Without Thinking
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
von: Gallego, Víctor
Veröffentlicht: (2025)
von: Gallego, Víctor
Veröffentlicht: (2025)
Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoning
von: Kansal, Yuval, et al.
Veröffentlicht: (2026)
von: Kansal, Yuval, et al.
Veröffentlicht: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
von: Ono, Shinnosuke, et al.
Veröffentlicht: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Process Reinforcement through Implicit Rewards
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
von: Wang, Ye, et al.
Veröffentlicht: (2026)
von: Wang, Ye, et al.
Veröffentlicht: (2026)
Negotiating with LLMS: Prompt Hacks, Skill Gaps, and Reasoning Deficits
von: Schneider, Johannes, et al.
Veröffentlicht: (2023)
von: Schneider, Johannes, et al.
Veröffentlicht: (2023)
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
von: Sheshanarayana, Disha, et al.
Veröffentlicht: (2026)
von: Sheshanarayana, Disha, et al.
Veröffentlicht: (2026)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
von: Kohli, Harsh, et al.
Veröffentlicht: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2026)
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
von: Liu, Runheng, et al.
Veröffentlicht: (2026)
von: Liu, Runheng, et al.
Veröffentlicht: (2026)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
von: Zhang, Zhen, et al.
Veröffentlicht: (2025)
Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
von: Yang, Wang, et al.
Veröffentlicht: (2026)
von: Yang, Wang, et al.
Veröffentlicht: (2026)
Brittleness and Promise: Knowledge Graph Based Reward Modeling for Diagnostic Reasoning
von: Khatwani, Saksham, et al.
Veröffentlicht: (2025)
von: Khatwani, Saksham, et al.
Veröffentlicht: (2025)
Towards Hierarchical Multi-Step Reward Models for Enhanced Reasoning in Large Language Models
von: Wang, Teng, et al.
Veröffentlicht: (2025)
von: Wang, Teng, et al.
Veröffentlicht: (2025)
Implicit Cross-Lingual Rewarding for Efficient Multilingual Preference Alignment
von: Yang, Wen, et al.
Veröffentlicht: (2025)
von: Yang, Wen, et al.
Veröffentlicht: (2025)
The Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning Models
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
von: Tan, Yuqiao, et al.
Veröffentlicht: (2025)
Detecting Prefix Bias in LLM-based Reward Models
von: Kumar, Ashwin, et al.
Veröffentlicht: (2025)
von: Kumar, Ashwin, et al.
Veröffentlicht: (2025)
Process Reward Models That Think
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
von: Khalifa, Muhammad, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
von: Si, Shengyun, et al.
Veröffentlicht: (2025) -
Spontaneous Reward Hacking in Iterative Self-Refinement
von: Pan, Jane, et al.
Veröffentlicht: (2024) -
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
von: Wang, Xinpeng, et al.
Veröffentlicht: (2024) -
LLMs Are Prone to Fallacies in Causal Inference
von: Joshi, Nitish, et al.
Veröffentlicht: (2024) -
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
von: Turpin, Miles, et al.
Veröffentlicht: (2025)