Resisting Correction: How RLHF Makes Language Models Ignore External Safety Signals in Natural Conversation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Cataneo, Felipe Biava |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models
von: Sun, Zhoujian, et al.
Veröffentlicht: (2024)
von: Sun, Zhoujian, et al.
Veröffentlicht: (2024)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
von: Yin, Yueqin, et al.
Veröffentlicht: (2025)
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
von: Zhou, Jiayi, et al.
Veröffentlicht: (2024)
von: Zhou, Jiayi, et al.
Veröffentlicht: (2024)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
von: Chen, Yanjun, et al.
Veröffentlicht: (2024)
Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation
von: Ling, Xinyi, et al.
Veröffentlicht: (2026)
von: Ling, Xinyi, et al.
Veröffentlicht: (2026)
Language Modeling with Editable External Knowledge
von: Li, Belinda Z., et al.
Veröffentlicht: (2024)
von: Li, Belinda Z., et al.
Veröffentlicht: (2024)
From RLHF to Direct Alignment: A Theoretical Unification of Preference Learning for Large Language Models
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
von: Raheja, Tarun, et al.
Veröffentlicht: (2026)
The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
von: Zhu, Yu, et al.
Veröffentlicht: (2024)
Automated Conversion of Static to Dynamic Scheduler via Natural Language
von: Tang, Paul Mingzheng, et al.
Veröffentlicht: (2024)
von: Tang, Paul Mingzheng, et al.
Veröffentlicht: (2024)
The Evolution of Natural Language Processing: How Prompt Optimization and Language Models are Shaping the Future
von: Saleem, Summra, et al.
Veröffentlicht: (2025)
von: Saleem, Summra, et al.
Veröffentlicht: (2025)
Reward Model Overoptimisation in Iterated RLHF
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
von: Wolf, Lorenz, et al.
Veröffentlicht: (2025)
Three Models of RLHF Annotation: Extension, Evidence, and Authority
von: Coyne, Steve
Veröffentlicht: (2026)
von: Coyne, Steve
Veröffentlicht: (2026)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
von: Noukhovitch, Michael, et al.
Veröffentlicht: (2024)
Retrieval Augmented Generation (RAG) and Beyond: A Comprehensive Survey on How to Make your LLMs use External Data More Wisely
von: Zhao, Siyun, et al.
Veröffentlicht: (2024)
von: Zhao, Siyun, et al.
Veröffentlicht: (2024)
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
von: Li, Zherui, et al.
Veröffentlicht: (2025)
von: Li, Zherui, et al.
Veröffentlicht: (2025)
Clarify: Improving Model Robustness With Natural Language Corrections
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
von: Lee, Yoonho, et al.
Veröffentlicht: (2024)
Prototypical Reward Network for Data-Efficient RLHF
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
von: Zhang, Jinghan, et al.
Veröffentlicht: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
von: Hu, Jian, et al.
Veröffentlicht: (2024)
von: Hu, Jian, et al.
Veröffentlicht: (2024)
Quantile Regression for Distributional Reward Models in RLHF
von: Dorka, Nicolai
Veröffentlicht: (2024)
von: Dorka, Nicolai
Veröffentlicht: (2024)
Identifying Multiple Personalities in Large Language Models with External Evaluation
von: Song, Xiaoyang, et al.
Veröffentlicht: (2024)
von: Song, Xiaoyang, et al.
Veröffentlicht: (2024)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
Reward Difference Optimization For Sample Reweighting In Offline RLHF
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
von: Wang, Shiqi, et al.
Veröffentlicht: (2024)
How to Retrieve Examples in In-context Learning to Improve Conversational Emotion Recognition using Large Language Models?
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
Deep Natural Language Feature Learning for Interpretable Prediction
von: Urrutia, Felipe, et al.
Veröffentlicht: (2023)
von: Urrutia, Felipe, et al.
Veröffentlicht: (2023)
Teaching Language Models How to Code Like Learners: Conversational Serialization for Student Simulation
von: Koutcheme, Charles, et al.
Veröffentlicht: (2026)
von: Koutcheme, Charles, et al.
Veröffentlicht: (2026)
Assessing Good, Bad and Ugly Arguments Generated by ChatGPT: a New Dataset, its Methodology and Associated Tasks
von: Rocha, Victor Hugo Nascimento, et al.
Veröffentlicht: (2024)
von: Rocha, Victor Hugo Nascimento, et al.
Veröffentlicht: (2024)
Rethinking External Slow-Thinking: From Snowball Errors to Probability of Correct Reasoning
von: Gan, Zeyu, et al.
Veröffentlicht: (2025)
von: Gan, Zeyu, et al.
Veröffentlicht: (2025)
Mitigating Length Bias in RLHF through a Causal Lens
von: Kim, Hyeonji, et al.
Veröffentlicht: (2025)
von: Kim, Hyeonji, et al.
Veröffentlicht: (2025)
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
von: Li, Aaron J., et al.
Veröffentlicht: (2024)
Removing RLHF Protections in GPT-4 via Fine-Tuning
von: Zhan, Qiusi, et al.
Veröffentlicht: (2023)
von: Zhan, Qiusi, et al.
Veröffentlicht: (2023)
RLHF and IIA: Perverse Incentives
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
von: Xu, Wanqiao, et al.
Veröffentlicht: (2023)
Reward-Robust RLHF in LLMs
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
von: Yan, Yuzi, et al.
Veröffentlicht: (2024)
Diversifying Question Generation over Knowledge Base via External Natural Questions
von: Guo, Shasha, et al.
Veröffentlicht: (2023)
von: Guo, Shasha, et al.
Veröffentlicht: (2023)
The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language Models
von: Zhang, Shuo, et al.
Veröffentlicht: (2023)
von: Zhang, Shuo, et al.
Veröffentlicht: (2023)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024) -
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models
von: Sun, Zhoujian, et al.
Veröffentlicht: (2024) -
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
von: Yin, Yueqin, et al.
Veröffentlicht: (2025) -
Sequence to Sequence Reward Modeling: Improving RLHF by Language Feedback
von: Zhou, Jiayi, et al.
Veröffentlicht: (2024) -
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)