Feedback Loops With Language Models Drive In-Context Reward Hacking
Fuente:
arXiv
Guardado en:
| Autores principales: | Pan, Alexander, Jones, Erik, Jagadeesan, Meena, Steinhardt, Jacob |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How do Language Models Bind Entities in Context?
por: Feng, Jiahai, et al.
Publicado: (2023)
por: Feng, Jiahai, et al.
Publicado: (2023)
Which Attention Heads Matter for In-Context Learning?
por: Yin, Kayo, et al.
Publicado: (2025)
por: Yin, Kayo, et al.
Publicado: (2025)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026)
por: Ackermann, Johannes, et al.
Publicado: (2026)
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025)
por: Fu, Jiayi, et al.
Publicado: (2025)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
por: Halawi, Danny, et al.
Publicado: (2023)
por: Halawi, Danny, et al.
Publicado: (2023)
Discovering Latent Knowledge in Language Models Without Supervision
por: Burns, Collin, et al.
Publicado: (2022)
por: Burns, Collin, et al.
Publicado: (2022)
Explaining Datasets in Words: Statistical Models with Natural Language Parameters
por: Zhong, Ruiqi, et al.
Publicado: (2024)
por: Zhong, Ruiqi, et al.
Publicado: (2024)
Training Language Models to Explain Their Own Computations
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
ODIN: Disentangled Reward Mitigates Hacking in RLHF
por: Chen, Lichang, et al.
Publicado: (2024)
por: Chen, Lichang, et al.
Publicado: (2024)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
por: Liu, Shang, et al.
Publicado: (2024)
por: Liu, Shang, et al.
Publicado: (2024)
On Teacher Hacking in Language Model Distillation
por: Tiapkin, Daniil, et al.
Publicado: (2025)
por: Tiapkin, Daniil, et al.
Publicado: (2025)
Mitigating Reward Hacking in RLHF via Advantage Sign Robustness
por: Ono, Shinnosuke, et al.
Publicado: (2026)
por: Ono, Shinnosuke, et al.
Publicado: (2026)
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
por: Ye, Yaowen, et al.
Publicado: (2025)
por: Ye, Yaowen, et al.
Publicado: (2025)
Approaching Human-Level Forecasting with Language Models
por: Halawi, Danny, et al.
Publicado: (2024)
por: Halawi, Danny, et al.
Publicado: (2024)
Temper and Tilt Lead to SLOP: Reward Hacking Mitigation with Inference-Time Alignment
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
Countdown-Code: A Testbed for Studying The Emergence and Generalization of Reward Hacking in RLVR
por: Khalifa, Muhammad, et al.
Publicado: (2026)
por: Khalifa, Muhammad, et al.
Publicado: (2026)
Learning a Generative Meta-Model of LLM Activations
por: Luo, Grace, et al.
Publicado: (2026)
por: Luo, Grace, et al.
Publicado: (2026)
Eliciting Language Model Behaviors with Investigator Agents
por: Li, Xiang Lisa, et al.
Publicado: (2025)
por: Li, Xiang Lisa, et al.
Publicado: (2025)
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
por: Jones, Erik, et al.
Publicado: (2025)
por: Jones, Erik, et al.
Publicado: (2025)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
por: Luo, Renjie, et al.
Publicado: (2025)
por: Luo, Renjie, et al.
Publicado: (2025)
Mass-Producing Failures of Multimodal Systems with Language Models
por: Tong, Shengbang, et al.
Publicado: (2023)
por: Tong, Shengbang, et al.
Publicado: (2023)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
por: Sahoo, Subramanyam
Publicado: (2026)
por: Sahoo, Subramanyam
Publicado: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
por: Wang, Hao, et al.
Publicado: (2026)
por: Wang, Hao, et al.
Publicado: (2026)
Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment
por: Yang, Rui, et al.
Publicado: (2024)
por: Yang, Rui, et al.
Publicado: (2024)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
por: Nair, Inderjeet, et al.
Publicado: (2024)
por: Nair, Inderjeet, et al.
Publicado: (2024)
Solve the Loop: Attractor Models for Language and Reasoning
por: Fein-Ashley, Jacob, et al.
Publicado: (2026)
por: Fein-Ashley, Jacob, et al.
Publicado: (2026)
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
por: Baumann, Joachim, et al.
Publicado: (2025)
por: Baumann, Joachim, et al.
Publicado: (2025)
Repairing Reward Functions with Feedback to Mitigate Reward Hacking
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
por: Hatgis-Kessell, Stephane, et al.
Publicado: (2025)
Online Intrinsic Rewards for Decision Making Agents from Large Language Model Feedback
por: Zheng, Qinqing, et al.
Publicado: (2024)
por: Zheng, Qinqing, et al.
Publicado: (2024)
Adversaries Can Misuse Combinations of Safe Models
por: Jones, Erik, et al.
Publicado: (2024)
por: Jones, Erik, et al.
Publicado: (2024)
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
por: Wang, Zhilin, et al.
Publicado: (2025)
por: Wang, Zhilin, et al.
Publicado: (2025)
Understanding In-context Learning of Addition via Activation Subspaces
por: Hu, Xinyan, et al.
Publicado: (2025)
por: Hu, Xinyan, et al.
Publicado: (2025)
Predictive Concept Decoders: Training Scalable End-to-End Interpretability Assistants
por: Huang, Vincent, et al.
Publicado: (2025)
por: Huang, Vincent, et al.
Publicado: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
por: Zhang, Shun, et al.
Publicado: (2024)
por: Zhang, Shun, et al.
Publicado: (2024)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
por: Ackermann, Johannes, et al.
Publicado: (2025)
por: Ackermann, Johannes, et al.
Publicado: (2025)
DynaSpec: Context-aware Dynamic Speculative Sampling for Large-Vocabulary Language Models
por: Zhang, Jinbin, et al.
Publicado: (2025)
por: Zhang, Jinbin, et al.
Publicado: (2025)
On the Robustness of Reward Models for Language Model Alignment
por: Hong, Jiwoo, et al.
Publicado: (2025)
por: Hong, Jiwoo, et al.
Publicado: (2025)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
por: Pan, Alexander, et al.
Publicado: (2024)
por: Pan, Alexander, et al.
Publicado: (2024)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
por: Cui, Ganqu, et al.
Publicado: (2023)
por: Cui, Ganqu, et al.
Publicado: (2023)
Reward Is Enough: LLMs Are In-Context Reinforcement Learners
por: Song, Kefan, et al.
Publicado: (2025)
por: Song, Kefan, et al.
Publicado: (2025)
Ejemplares similares
-
How do Language Models Bind Entities in Context?
por: Feng, Jiahai, et al.
Publicado: (2023) -
Which Attention Heads Matter for In-Context Learning?
por: Yin, Kayo, et al.
Publicado: (2025) -
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
por: Ackermann, Johannes, et al.
Publicado: (2026) -
Reward Shaping to Mitigate Reward Hacking in RLHF
por: Fu, Jiayi, et al.
Publicado: (2025) -
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
por: Halawi, Danny, et al.
Publicado: (2023)