When Helpfulness Becomes Sycophancy: Sycophancy is a Boundary Failure Between Social Alignment and Epistemic Integrity in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jiechen, Barry, Catherine A., Randev, Rishika, Chen, Janet, Jorgensen, Ella, Bent, Brinnae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026)
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
von: Bent, Brinnae
Veröffentlicht: (2024)
von: Bent, Brinnae
Veröffentlicht: (2024)
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024)
von: Malmqvist, Lars
Veröffentlicht: (2024)
Consistency Training Helps Stop Sycophancy and Jailbreaks
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
von: Bent, Brinnae
Veröffentlicht: (2025)
von: Bent, Brinnae
Veröffentlicht: (2025)
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
von: Wang, Keyu, et al.
Veröffentlicht: (2025)
Towards Understanding Sycophancy in Language Models
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
Moral Sycophancy in Vision Language Models
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models
von: Nama, Vihaan, et al.
Veröffentlicht: (2026)
von: Nama, Vihaan, et al.
Veröffentlicht: (2026)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
von: Denison, Carson, et al.
Veröffentlicht: (2024)
von: Denison, Carson, et al.
Veröffentlicht: (2024)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
von: Guo, Kevin H., et al.
Veröffentlicht: (2026)
Spatiotemporal Sycophancy: Negation-Based Gaslighting in Video Large Language Models
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
von: Tang, Ziyao, et al.
Veröffentlicht: (2026)
Accounting for Sycophancy in Language Model Uncertainty Estimation
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
von: Kasneci, Enkelejda, et al.
Veröffentlicht: (2026)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2025)
GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
von: Batzner, Jan, et al.
Veröffentlicht: (2024)
von: Batzner, Jan, et al.
Veröffentlicht: (2024)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
von: Hong, Jiseung, et al.
Veröffentlicht: (2025)
Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
von: Xu, Juangui, et al.
Veröffentlicht: (2025)
It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt
von: Bladon, Stuart, et al.
Veröffentlicht: (2026)
von: Bladon, Stuart, et al.
Veröffentlicht: (2026)
Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2
von: Bladon, Stuart, et al.
Veröffentlicht: (2026)
von: Bladon, Stuart, et al.
Veröffentlicht: (2026)
Human-like Social Compliance in Large Language Models: Unifying Sycophancy and Conformity through Signal Competition Dynamics
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
How RLHF Amplifies Sycophancy
von: Shapira, Itai, et al.
Veröffentlicht: (2026)
von: Shapira, Itai, et al.
Veröffentlicht: (2026)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
von: Pandey, Sanskar, et al.
Veröffentlicht: (2025)
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
von: Batzner, Jan, et al.
Veröffentlicht: (2025)
von: Batzner, Jan, et al.
Veröffentlicht: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
von: Liu, Joshua, et al.
Veröffentlicht: (2025)
von: Liu, Joshua, et al.
Veröffentlicht: (2025)
Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
von: Christian, Brian, et al.
Veröffentlicht: (2026)
von: Christian, Brian, et al.
Veröffentlicht: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
von: Atwell, Katherine, et al.
Veröffentlicht: (2025)
Sycophancy Hides Linearly in the Attention Heads
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
von: Genadi, Rifo, et al.
Veröffentlicht: (2026)
Sycophancy as compositions of Atomic Psychometric Traits
von: Jain, Shreyans, et al.
Veröffentlicht: (2025)
von: Jain, Shreyans, et al.
Veröffentlicht: (2025)
SycEval: Evaluating LLM Sycophancy
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
von: Fanous, Aaron, et al.
Veröffentlicht: (2025)
From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
The Social Sycophancy Scale: A psychometrically validated measure of sycophancy
von: Rehani, Jean, et al.
Veröffentlicht: (2026)
von: Rehani, Jean, et al.
Veröffentlicht: (2026)
To Agree or To Be Right? The Grounding-Sycophancy Tradeoff in Medical Vision-Language Models
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
Hearing is Believing? Evaluating and Analyzing Audio Language Model Sycophancy with SYAUDIO
von: Yao, Junchi, et al.
Veröffentlicht: (2026)
von: Yao, Junchi, et al.
Veröffentlicht: (2026)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
Intersectional Sycophancy: How Perceived User Demographics Shape False Validation in Large Language Models
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
von: Maltbie, Benjamin, et al.
Veröffentlicht: (2026)
Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models
von: Chang, Edward Y.
Veröffentlicht: (2025)
von: Chang, Edward Y.
Veröffentlicht: (2025)
The Illusion of Agreement with ChatGPT: Sycophancy and Beyond
von: Noshin, Kazi, et al.
Veröffentlicht: (2026)
von: Noshin, Kazi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
von: Natan, Shahar Ben, et al.
Veröffentlicht: (2026) -
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
von: Bent, Brinnae
Veröffentlicht: (2024) -
Sycophancy in Large Language Models: Causes and Mitigations
von: Malmqvist, Lars
Veröffentlicht: (2024) -
Consistency Training Helps Stop Sycophancy and Jailbreaks
von: Irpan, Alex, et al.
Veröffentlicht: (2025) -
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
von: Bent, Brinnae
Veröffentlicht: (2025)