Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Fuente:
arXiv
Salvato in:
| Autori principali: | Ferreira, Pedro, Aziz, Wilker, Titov, Ivan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Explanation Regularisation through the Lens of Attributions
di: Ferreira, Pedro, et al.
Pubblicazione: (2024)
di: Ferreira, Pedro, et al.
Pubblicazione: (2024)
RRM: Robust Reward Model Training Mitigates Reward Hacking
di: Liu, Tianqi, et al.
Pubblicazione: (2024)
di: Liu, Tianqi, et al.
Pubblicazione: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
Predict the Next Word: Humans exhibit uncertainty in this task and language models _____
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
di: Ilia, Evgenia, et al.
Pubblicazione: (2024)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
di: Liu, Yifeng, et al.
Pubblicazione: (2026)
di: Liu, Yifeng, et al.
Pubblicazione: (2026)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
di: Wu, Rui, et al.
Pubblicazione: (2026)
di: Wu, Rui, et al.
Pubblicazione: (2026)
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models
di: Deng, Wenlong, et al.
Pubblicazione: (2026)
di: Deng, Wenlong, et al.
Pubblicazione: (2026)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
MedPath: Multi-Domain Cross-Vocabulary Hierarchical Paths for Biomedical Entity Linking
di: Mishra, Nishant, et al.
Pubblicazione: (2025)
di: Mishra, Nishant, et al.
Pubblicazione: (2025)
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
di: Gallego, Víctor
Pubblicazione: (2025)
di: Gallego, Víctor
Pubblicazione: (2025)
Mitigating Copy Bias in In-Context Learning through Neuron Pruning
di: Ali, Ameen, et al.
Pubblicazione: (2024)
di: Ali, Ameen, et al.
Pubblicazione: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
di: Wang, Ye, et al.
Pubblicazione: (2026)
di: Wang, Ye, et al.
Pubblicazione: (2026)
Belief Attribution as Mental Explanation: The Role of Accuracy, Informativity, and Causality
di: Ying, Lance, et al.
Pubblicazione: (2025)
di: Ying, Lance, et al.
Pubblicazione: (2025)
Interpreting Predictive Probabilities: Model Confidence or Human Label Variation?
di: Baan, Joris, et al.
Pubblicazione: (2024)
di: Baan, Joris, et al.
Pubblicazione: (2024)
Spontaneous Reward Hacking in Iterative Self-Refinement
di: Pan, Jane, et al.
Pubblicazione: (2024)
di: Pan, Jane, et al.
Pubblicazione: (2024)
Detecting and Suppressing Reward Hacking with Gradient Fingerprints
di: Wang, Songtao, et al.
Pubblicazione: (2026)
di: Wang, Songtao, et al.
Pubblicazione: (2026)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
Generalisation First, Memorisation Second? Memorisation Localisation for Natural Language Classification Tasks
di: Dankers, Verna, et al.
Pubblicazione: (2024)
di: Dankers, Verna, et al.
Pubblicazione: (2024)
Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
di: Zang, Jianxiang
Pubblicazione: (2025)
di: Zang, Jianxiang
Pubblicazione: (2025)
Post-hoc Reward Calibration: A Case Study on Length Bias
di: Huang, Zeyu, et al.
Pubblicazione: (2024)
di: Huang, Zeyu, et al.
Pubblicazione: (2024)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
di: Turpin, Miles, et al.
Pubblicazione: (2025)
di: Turpin, Miles, et al.
Pubblicazione: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
di: Baan, Joris, et al.
Pubblicazione: (2026)
di: Baan, Joris, et al.
Pubblicazione: (2026)
M-Wanda: Improving One-Shot Pruning for Multilingual LLMs
di: Choenni, Rochelle, et al.
Pubblicazione: (2025)
di: Choenni, Rochelle, et al.
Pubblicazione: (2025)
Unlearning Traces the Influential Training Data of Language Models
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
di: Wang, Xinpeng, et al.
Pubblicazione: (2025)
di: Wang, Xinpeng, et al.
Pubblicazione: (2025)
Monitoring Emergent Reward Hacking During Generation via Internal Activations
di: Wilhelm, Patrick, et al.
Pubblicazione: (2026)
di: Wilhelm, Patrick, et al.
Pubblicazione: (2026)
Feedback Loops With Language Models Drive In-Context Reward Hacking
di: Pan, Alexander, et al.
Pubblicazione: (2024)
di: Pan, Alexander, et al.
Pubblicazione: (2024)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
di: Alon, Bar, et al.
Pubblicazione: (2026)
di: Alon, Bar, et al.
Pubblicazione: (2026)
Tailored Truths: Optimizing LLM Persuasion with Personalization and Fabricated Statistics
di: Timm, Jasper, et al.
Pubblicazione: (2025)
di: Timm, Jasper, et al.
Pubblicazione: (2025)
Optimising Calls to Large Language Models with Uncertainty-Based Two-Tier Selection
di: Ramírez, Guillem, et al.
Pubblicazione: (2024)
di: Ramírez, Guillem, et al.
Pubblicazione: (2024)
SIP: Injecting a Structural Inductive Bias into a Seq2Seq Model by Simulation
di: Lindemann, Matthias, et al.
Pubblicazione: (2023)
di: Lindemann, Matthias, et al.
Pubblicazione: (2023)
Strengthening Structural Inductive Biases by Pre-training to Perform Syntactic Transformations
di: Lindemann, Matthias, et al.
Pubblicazione: (2024)
di: Lindemann, Matthias, et al.
Pubblicazione: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
What's New in My Data? Novelty Exploration via Contrastive Generation
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
di: Isonuma, Masaru, et al.
Pubblicazione: (2024)
Teaching Language Models to Faithfully Express their Uncertainty
di: Eikema, Bryan, et al.
Pubblicazione: (2025)
di: Eikema, Bryan, et al.
Pubblicazione: (2025)
Reward Engineering for Generating Semi-structured Explanation
di: Han, Jiuzhou, et al.
Pubblicazione: (2023)
di: Han, Jiuzhou, et al.
Pubblicazione: (2023)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
The Facade of Truth: Uncovering and Mitigating LLM Susceptibility to Deceptive Evidence
di: Wan, Herun, et al.
Pubblicazione: (2026)
di: Wan, Herun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Explanation Regularisation through the Lens of Attributions
di: Ferreira, Pedro, et al.
Pubblicazione: (2024) -
RRM: Robust Reward Model Training Mitigates Reward Hacking
di: Liu, Tianqi, et al.
Pubblicazione: (2024) -
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025) -
Variability Need Not Imply Error: The Case of Adequate but Semantically Distinct Responses
di: Ilia, Evgenia, et al.
Pubblicazione: (2024) -
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)