How RLHF Amplifies Sycophancy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shapira, Itai, Benade, Gerdus, Procaccia, Ariel D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pairwise Calibrated Rewards for Pluralistic Alignment
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
von: Halpern, Daniel, et al.
Veröffentlicht: (2025)
Axioms for AI Alignment from Human Feedback
von: Ge, Luise, et al.
Veröffentlicht: (2024)
von: Ge, Luise, et al.
Veröffentlicht: (2024)
Generative Social Choice
von: Fish, Sara, et al.
Veröffentlicht: (2023)
von: Fish, Sara, et al.
Veröffentlicht: (2023)
Offline Local Search for Online Stochastic Bandits
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026)
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026)
Incentives in Federated Learning with Heterogeneous Agents
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
Embeddings for Preferences, Not Semantics
von: Blair, Carter, et al.
Veröffentlicht: (2026)
von: Blair, Carter, et al.
Veröffentlicht: (2026)
Generative Social Choice: The Next Generation
von: Boehmer, Niclas, et al.
Veröffentlicht: (2025)
von: Boehmer, Niclas, et al.
Veröffentlicht: (2025)
Clone-Robust AI Alignment
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)
Learning Social Welfare Functions
von: Pardeshi, Kanad Shrikar, et al.
Veröffentlicht: (2024)
von: Pardeshi, Kanad Shrikar, et al.
Veröffentlicht: (2024)
Policy Aggregation
von: Alamdari, Parand A., et al.
Veröffentlicht: (2024)
von: Alamdari, Parand A., et al.
Veröffentlicht: (2024)
Moral Sycophancy in Vision Language Models
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
SycEval: Evaluating LLM Sycophancy
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
How to Evaluate Reward Models for RLHF
von: Frick, Evan, et al.
Veröffentlicht: (2024)
von: Frick, Evan, et al.
Veröffentlicht: (2024)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
Bias Detection Via Signaling
von: Chen, Yiling, et al.
Veröffentlicht: (2024)
von: Chen, Yiling, et al.
Veröffentlicht: (2024)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
Linear Probe Penalties Reduce LLM Sycophancy
von: Papadatos, Henry, et al.
Veröffentlicht: (2024)
von: Papadatos, Henry, et al.
Veröffentlicht: (2024)
When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
von: Li, Jiechen, et al.
Veröffentlicht: (2026)
von: Li, Jiechen, et al.
Veröffentlicht: (2026)
Alternates, Assemble! Selecting Optimal Alternates for Citizens' Assemblies
von: Assos, Angelos, et al.
Veröffentlicht: (2025)
von: Assos, Angelos, et al.
Veröffentlicht: (2025)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
von: Törnberg, Petter, et al.
Veröffentlicht: (2026)
von: Törnberg, Petter, et al.
Veröffentlicht: (2026)
Strategic Classification With Externalities
von: Hossain, Safwan, et al.
Veröffentlicht: (2024)
von: Hossain, Safwan, et al.
Veröffentlicht: (2024)
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024)
von: Malmqvist, Lars
Veröffentlicht: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
Question the Questions: Auditing Representation in Online Deliberative Processes
von: De, Soham, et al.
Veröffentlicht: (2025)
von: De, Soham, et al.
Veröffentlicht: (2025)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
Finding Common Ground in a Sea of Alternatives
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
von: Chooi, Jay, et al.
Veröffentlicht: (2026)
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
von: Wang, Libo
Veröffentlicht: (2024)
von: Wang, Libo
Veröffentlicht: (2024)
PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
von: Dong, Hanze, et al.
Veröffentlicht: (2024)
Direct Alignment with Heterogeneous Preferences
von: Shirali, Ali, et al.
Veröffentlicht: (2025)
von: Shirali, Ali, et al.
Veröffentlicht: (2025)
Pressure, What Pressure? Sycophancy Disentanglement in Language Models via Reward Decomposition
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
von: Mohsin, Muhammad Ahmed, et al.
Veröffentlicht: (2026)
Sycophancy Mitigation Through Reinforcement Learning with Uncertainty-Aware Adaptive Reasoning Trajectories
von: Beigi, Mohammad, et al.
Veröffentlicht: (2025)
von: Beigi, Mohammad, et al.
Veröffentlicht: (2025)
Mitigating Cognitive Bias in RLHF by Altering Rationality
von: Horter, Tiffany, et al.
Veröffentlicht: (2026)
von: Horter, Tiffany, et al.
Veröffentlicht: (2026)
Consistency Amplifies: How Behavioral Variance Shapes Agent Accuracy
von: Mehta, Aman
Veröffentlicht: (2026)
von: Mehta, Aman
Veröffentlicht: (2026)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
ROCM: RLHF on consistency models
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2025)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pairwise Calibrated Rewards for Pluralistic Alignment
von: Halpern, Daniel, et al.
Veröffentlicht: (2025) -
Axioms for AI Alignment from Human Feedback
von: Ge, Luise, et al.
Veröffentlicht: (2024) -
Generative Social Choice
von: Fish, Sara, et al.
Veröffentlicht: (2023) -
Offline Local Search for Online Stochastic Bandits
von: Benadè, Gerdus, et al.
Veröffentlicht: (2026) -
Incentives in Federated Learning with Heterogeneous Agents
von: Procaccia, Ariel D., et al.
Veröffentlicht: (2025)