Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hahm, Dongyoon, Hadfield-Menell, Dylan, Lee, Kimin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
Enhancing LLM Agent Safety via Causal Influence Prompting
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025)
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
von: Ma, Rachel, et al.
Veröffentlicht: (2025)
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023)
Benchmarking Mobile Device Control Agents across Diverse Configurations
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Disjoint Processing Mechanisms of Hierarchical and Linear Grammars in Large Language Models
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2025)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
von: Lee, Juyong, et al.
Veröffentlicht: (2024)
Distributional Process Reward Models: Calibrated Prediction of Future Rewards via Conditional Optimal Transport
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
von: Ma, Rachel, et al.
Veröffentlicht: (2026)
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
von: Lee, Harrison, et al.
Veröffentlicht: (2023)
Distributionally Robust Reinforcement Learning with Human Feedback
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
von: Mandal, Debmalya, et al.
Veröffentlicht: (2025)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
von: Giordani, Jeremiah
Veröffentlicht: (2025)
von: Giordani, Jeremiah
Veröffentlicht: (2025)
Prompt Injection as Role Confusion
von: Ye, Charles, et al.
Veröffentlicht: (2026)
von: Ye, Charles, et al.
Veröffentlicht: (2026)
Parameter Efficient Reinforcement Learning from Human Feedback
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
von: Sidahmed, Hakim, et al.
Veröffentlicht: (2024)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
RLTHF: Targeted Human Feedback for LLM Alignment
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
von: Xu, Yifei, et al.
Veröffentlicht: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Understanding Impact of Human Feedback via Influence Functions
von: Min, Taywon, et al.
Veröffentlicht: (2025)
von: Min, Taywon, et al.
Veröffentlicht: (2025)
On the Hidden Objective Biases of Group-based Reinforcement Learning
von: Fontana, Aleksandar, et al.
Veröffentlicht: (2026)
von: Fontana, Aleksandar, et al.
Veröffentlicht: (2026)
Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2025)
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
von: Zhang, Shun, et al.
Veröffentlicht: (2024)
von: Zhang, Shun, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Backtracking Feedback
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning
von: Poddar, Sriyash, et al.
Veröffentlicht: (2024)
von: Poddar, Sriyash, et al.
Veröffentlicht: (2024)
Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
von: Ma, Rachel, et al.
Veröffentlicht: (2024)
Tamper-Resistant Safeguards for Open-Weight LLMs
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
von: Tamirisa, Rishub, et al.
Veröffentlicht: (2024)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
von: Chaudhari, Shreyas, et al.
Veröffentlicht: (2024)
Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
Reinforcement Learning from Human Feedback with Active Queries
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
von: Ji, Kaixuan, et al.
Veröffentlicht: (2024)
Gradient Regularization Prevents Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
von: Ackermann, Johannes, et al.
Veröffentlicht: (2026)
Defending Against Unforeseen Failure Modes with Latent Adversarial Training
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
von: Casper, Stephen, et al.
Veröffentlicht: (2024)
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxing, et al.
Veröffentlicht: (2024)
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
von: Li, Gen, et al.
Veröffentlicht: (2025)
von: Li, Gen, et al.
Veröffentlicht: (2025)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
von: Huang, Audrey, et al.
Veröffentlicht: (2024)
Large Language Models are Biased Reinforcement Learners
von: Hayes, William M., et al.
Veröffentlicht: (2024)
von: Hayes, William M., et al.
Veröffentlicht: (2024)
Learning Personalized Agents from Human Feedback
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2026)
Can We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning Models
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
von: Chan, Yik Siu, et al.
Veröffentlicht: (2025)
Activation Steering via Generative Causal Mediation
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
von: Sankaranarayanan, Aruna, et al.
Veröffentlicht: (2026)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
von: Cheng, Ruoxi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025) -
Enhancing LLM Agent Safety via Causal Influence Prompting
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2025) -
Flexible Agent Alignment with Goal Inference from Open-Ended Dialog
von: Ma, Rachel, et al.
Veröffentlicht: (2025) -
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF
von: Siththaranjan, Anand, et al.
Veröffentlicht: (2023) -
Benchmarking Mobile Device Control Agents across Diverse Configurations
von: Lee, Juyong, et al.
Veröffentlicht: (2024)