Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Vennemeyer, Daniel, Duong, Phan Anh, Zhan, Tiffany, Jiang, Tianyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026)
by: Vennemeyer, Daniel, et al.
Published: (2026)
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
by: Hutson, Dylan, et al.
Published: (2025)
by: Hutson, Dylan, et al.
Published: (2025)
Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs
by: Aldahlawi, Bayan Abdullah, et al.
Published: (2026)
by: Aldahlawi, Bayan Abdullah, et al.
Published: (2026)
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
by: Zhang, Kaiwei, et al.
Published: (2025)
by: Zhang, Kaiwei, et al.
Published: (2025)
CHEER-Ekman: Fine-grained Embodied Emotion Classification
by: Duong, Phan Anh, et al.
Published: (2025)
by: Duong, Phan Anh, et al.
Published: (2025)
Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
by: Saim, Mohammad, et al.
Published: (2025)
by: Saim, Mohammad, et al.
Published: (2025)
Overalignment in Frontier LLMs: An Empirical Study of Sycophantic Behaviour in Healthcare
by: Christophe, Clément, et al.
Published: (2026)
by: Christophe, Clément, et al.
Published: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)
by: Zhou, Wenrui, et al.
Published: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
by: Natan, Shahar Ben, et al.
Published: (2026)
by: Natan, Shahar Ben, et al.
Published: (2026)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
by: Zhang, Chaowei, et al.
Published: (2026)
by: Zhang, Chaowei, et al.
Published: (2026)
Hierarchical Document Parsing via Large Margin Feature Matching and Heuristics
by: Kiet, Duong Anh
Published: (2025)
by: Kiet, Duong Anh
Published: (2025)
When AI Tells You What You Want to Hear: Sycophantic Behavior of Large Language Models in Dementia Care Settings
by: Kolb, Christian
Published: (2026)
by: Kolb, Christian
Published: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
by: Yao, Louie Hong, et al.
Published: (2025)
by: Yao, Louie Hong, et al.
Published: (2025)
FVA-RAG: Falsification-Verification Alignment for Mitigating Sycophantic Hallucinations
by: Ravishankara, Mayank
Published: (2025)
by: Ravishankara, Mayank
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
by: Hong, Jiseung, et al.
Published: (2025)
by: Hong, Jiseung, et al.
Published: (2025)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
by: Do, Heejin, et al.
Published: (2026)
by: Do, Heejin, et al.
Published: (2026)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
by: Yao, Binwei, et al.
Published: (2025)
by: Yao, Binwei, et al.
Published: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
by: Liu, Joshua, et al.
Published: (2025)
by: Liu, Joshua, et al.
Published: (2025)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
by: Nogueira, Rodrigo, et al.
Published: (2026)
by: Nogueira, Rodrigo, et al.
Published: (2026)
When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
by: Ranaldi, Leonardo, et al.
Published: (2023)
by: Ranaldi, Leonardo, et al.
Published: (2023)
Sycophancy in Large Language Models: Causes and Mitigations
by: Malmqvist, Lars
Published: (2024)
by: Malmqvist, Lars
Published: (2024)
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
by: Batzner, Jan, et al.
Published: (2025)
by: Batzner, Jan, et al.
Published: (2025)
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
by: Arvin, Chuck
Published: (2025)
by: Arvin, Chuck
Published: (2025)
SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
by: Bhalla, Joy, et al.
Published: (2026)
by: Bhalla, Joy, et al.
Published: (2026)
Large Language Models Often Say One Thing and Do Another
by: Xu, Ruoxi, et al.
Published: (2025)
by: Xu, Ruoxi, et al.
Published: (2025)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
by: Çelebi, Yusuf, et al.
Published: (2025)
by: Çelebi, Yusuf, et al.
Published: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
by: Wang, Keyu, et al.
Published: (2025)
by: Wang, Keyu, et al.
Published: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
Do LLMs Encode Frame Semantics? Evidence from Frame Identification
by: Chundru, Jayanth Krishna, et al.
Published: (2025)
by: Chundru, Jayanth Krishna, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
Similar Items
-
Objective Matters: Fine-Tuning Objectives Shape Safety, Robustness, and Persona Drift
by: Vennemeyer, Daniel, et al.
Published: (2026) -
GuessingGame: Measuring the Informativeness of Open-Ended Questions in Large Language Models
by: Hutson, Dylan, et al.
Published: (2025) -
Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs
by: Aldahlawi, Bayan Abdullah, et al.
Published: (2026) -
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
by: Zhang, Kaiwei, et al.
Published: (2025) -
CHEER-Ekman: Fine-grained Embodied Emotion Classification
by: Duong, Phan Anh, et al.
Published: (2025)