Spontaneous Reward Hacking in Iterative Self-Refinement
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Jane, He, He, Bowman, Samuel R., Feng, Shi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
di: Gallego, Víctor
Pubblicazione: (2025)
di: Gallego, Víctor
Pubblicazione: (2025)
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
di: Wang, Xinpeng, et al.
Pubblicazione: (2025)
di: Wang, Xinpeng, et al.
Pubblicazione: (2025)
LLM Evaluators Recognize and Favor Their Own Generations
di: Panickssery, Arjun, et al.
Pubblicazione: (2024)
di: Panickssery, Arjun, et al.
Pubblicazione: (2024)
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
di: Fu, Jiayi, et al.
Pubblicazione: (2025)
Feedback Loops With Language Models Drive In-Context Reward Hacking
di: Pan, Alexander, et al.
Pubblicazione: (2024)
di: Pan, Alexander, et al.
Pubblicazione: (2024)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
di: Turpin, Miles, et al.
Pubblicazione: (2025)
di: Turpin, Miles, et al.
Pubblicazione: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
di: Chen, Lichang, et al.
Pubblicazione: (2024)
di: Chen, Lichang, et al.
Pubblicazione: (2024)
Monitoring Emergent Reward Hacking During Generation via Internal Activations
di: Wilhelm, Patrick, et al.
Pubblicazione: (2026)
di: Wilhelm, Patrick, et al.
Pubblicazione: (2026)
AIR: Complex Instruction Generation via Automatic Iterative Refinement
di: Liu, Wei, et al.
Pubblicazione: (2025)
di: Liu, Wei, et al.
Pubblicazione: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
di: Ono, Shinnosuke, et al.
Pubblicazione: (2026)
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification
di: Zhang, Anqi, et al.
Pubblicazione: (2025)
di: Zhang, Anqi, et al.
Pubblicazione: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
di: Ackermann, Johannes, et al.
Pubblicazione: (2026)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
di: Zhao, Bingchen, et al.
Pubblicazione: (2026)
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
di: Liu, Xiaoyuan, et al.
Pubblicazione: (2025)
SCIR: A Self-Correcting Iterative Refinement Framework for Enhanced Information Extraction Based on Schema
di: Fang, Yushen, et al.
Pubblicazione: (2025)
di: Fang, Yushen, et al.
Pubblicazione: (2025)
Iterative Translation Refinement with Large Language Models
di: Chen, Pinzhen, et al.
Pubblicazione: (2023)
di: Chen, Pinzhen, et al.
Pubblicazione: (2023)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
di: Wang, Ye, et al.
Pubblicazione: (2026)
di: Wang, Ye, et al.
Pubblicazione: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
di: Khalifa, Muhammad, et al.
Pubblicazione: (2026)
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
di: Xu, Wenda, et al.
Pubblicazione: (2024)
di: Xu, Wenda, et al.
Pubblicazione: (2024)
Iterative Reasoning Preference Optimization
di: Pang, Richard Yuanzhe, et al.
Pubblicazione: (2024)
di: Pang, Richard Yuanzhe, et al.
Pubblicazione: (2024)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
di: Denison, Carson, et al.
Pubblicazione: (2024)
di: Denison, Carson, et al.
Pubblicazione: (2024)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
di: Padmakumar, Vishakh, et al.
Pubblicazione: (2024)
di: Padmakumar, Vishakh, et al.
Pubblicazione: (2024)
De Jure: Iterative LLM Self-Refinement for Structured Extraction of Regulatory Rules
di: Guliani, Keerat, et al.
Pubblicazione: (2026)
di: Guliani, Keerat, et al.
Pubblicazione: (2026)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
di: Yu, Simon, et al.
Pubblicazione: (2024)
di: Yu, Simon, et al.
Pubblicazione: (2024)
Tiny Reward Models
di: Pan, Sarah
Pubblicazione: (2025)
di: Pan, Sarah
Pubblicazione: (2025)
RefineCoder: Iterative Improving of Large Language Models via Adaptive Critique Refinement for Code Generation
di: Zhou, Changzhi, et al.
Pubblicazione: (2025)
di: Zhou, Changzhi, et al.
Pubblicazione: (2025)
Reward Model Overoptimisation in Iterated RLHF
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
di: Wolf, Lorenz, et al.
Pubblicazione: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
di: Sahoo, Subramanyam
Pubblicazione: (2026)
di: Sahoo, Subramanyam
Pubblicazione: (2026)
Self-Rewarding Language Models
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
di: Yuan, Weizhe, et al.
Pubblicazione: (2024)
SEG:Seeds-Enhanced Iterative Refinement Graph Neural Network for Entity Alignment
di: Ai, Wei, et al.
Pubblicazione: (2024)
di: Ai, Wei, et al.
Pubblicazione: (2024)
FLAIRR-TS -- Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Series
di: Jalori, Gunjan, et al.
Pubblicazione: (2025)
di: Jalori, Gunjan, et al.
Pubblicazione: (2025)
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2024)
di: Liu, Chris Yuhao, et al.
Pubblicazione: (2024)
History-Guided Iterative Visual Reasoning with Self-Correction
di: Yang, Xinglong, et al.
Pubblicazione: (2026)
di: Yang, Xinglong, et al.
Pubblicazione: (2026)
TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
di: Feng, Zhaopeng, et al.
Pubblicazione: (2024)
di: Feng, Zhaopeng, et al.
Pubblicazione: (2024)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
di: Gao, Jiaxuan, et al.
Pubblicazione: (2026)
di: Gao, Jiaxuan, et al.
Pubblicazione: (2026)
Improving Machine Translation with Human Feedback: An Exploration of Quality Estimation as a Reward Model
di: He, Zhiwei, et al.
Pubblicazione: (2024)
di: He, Zhiwei, et al.
Pubblicazione: (2024)
Self-Evolved Reward Learning for LLMs
di: Huang, Chenghua, et al.
Pubblicazione: (2024)
di: Huang, Chenghua, et al.
Pubblicazione: (2024)
Self-Improvement as Coherence Optimization: A Theoretical Account
di: Qiu, Tianyi, et al.
Pubblicazione: (2026)
di: Qiu, Tianyi, et al.
Pubblicazione: (2026)
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
di: C, Rajmohan, et al.
Pubblicazione: (2025)
di: C, Rajmohan, et al.
Pubblicazione: (2025)
Iterative Prompt Refinement for Dyslexia-Friendly Text Summarization Using GPT-4o
di: Bhojwani, Samay, et al.
Pubblicazione: (2026)
di: Bhojwani, Samay, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Specification Self-Correction: Mitigating In-Context Reward Hacking Through Test-Time Refinement
di: Gallego, Víctor
Pubblicazione: (2025) -
Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
di: Wang, Xinpeng, et al.
Pubblicazione: (2025) -
LLM Evaluators Recognize and Favor Their Own Generations
di: Panickssery, Arjun, et al.
Pubblicazione: (2024) -
Reward Shaping to Mitigate Reward Hacking in RLHF
di: Fu, Jiayi, et al.
Pubblicazione: (2025) -
Feedback Loops With Language Models Drive In-Context Reward Hacking
di: Pan, Alexander, et al.
Pubblicazione: (2024)