LLMs Know They're Wrong and Agree Anyway: The Shared Sycophancy-Lying Circuit
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Pandey, Manav |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Androids Know They're Only Dreaming of Electric Sheep?
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023)
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
von: Xiaohu, Xie, et al.
Veröffentlicht: (2026)
von: Xiaohu, Xie, et al.
Veröffentlicht: (2026)
AgreeMate: Teaching LLMs to Haggle
von: Chatterjee, Ainesh, et al.
Veröffentlicht: (2024)
von: Chatterjee, Ainesh, et al.
Veröffentlicht: (2024)
Letting Others Know How They're Doing.
von: Hartzell, Gary
Veröffentlicht: (1993)
von: Hartzell, Gary
Veröffentlicht: (1993)
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
von: Magid, Salma Abdel, et al.
Veröffentlicht: (2024)
von: Magid, Salma Abdel, et al.
Veröffentlicht: (2024)
Can AI Agents Agree?
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
von: Berdoz, Frédéric, et al.
Veröffentlicht: (2026)
Do Two AI Scientists Agree?
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
von: Petrov, Ivo, et al.
Veröffentlicht: (2025)
Do Small Language Models Know When They're Wrong? Confidence-Based Cascade Scoring for Educational Assessment
von: Burleigh, Tyler
Veröffentlicht: (2026)
von: Burleigh, Tyler
Veröffentlicht: (2026)
Easy Problems That LLMs Get Wrong
von: Williams, Sean, et al.
Veröffentlicht: (2024)
von: Williams, Sean, et al.
Veröffentlicht: (2024)
They're Back! Invite Them In!
von: Barron, Daniel D.
Veröffentlicht: (1992)
von: Barron, Daniel D.
Veröffentlicht: (1992)
A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
von: O'Brien, Claire, et al.
Veröffentlicht: (2026)
von: O'Brien, Claire, et al.
Veröffentlicht: (2026)
Consistency Training Helps Stop Sycophancy and Jailbreaks
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
von: Irpan, Alex, et al.
Veröffentlicht: (2025)
LLMs Know When They Know, but Do Not Act on It: A Metacognitive Harness for Test-time Scaling
von: Cao, Qi, et al.
Veröffentlicht: (2026)
von: Cao, Qi, et al.
Veröffentlicht: (2026)
"Patriarchy Hurts Men Too." Does Your Model Agree? A Discussion on Fairness Assumptions
von: Favier, Marco, et al.
Veröffentlicht: (2024)
von: Favier, Marco, et al.
Veröffentlicht: (2024)
Do Language Models Know When They're Hallucinating References?
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
Bayesian Mixture-of-Experts: Towards Making LLMs Know What They Don't Know
von: Li, Albus Yizhuo
Veröffentlicht: (2025)
von: Li, Albus Yizhuo
Veröffentlicht: (2025)
Adapt, Agree, Aggregate: Semi-Supervised Ensemble Labeling for Graph Convolutional Networks
von: Abdolali, Maryam, et al.
Veröffentlicht: (2025)
von: Abdolali, Maryam, et al.
Veröffentlicht: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
von: Sahoo, Subramanyam
Veröffentlicht: (2026)
The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
Europe. They're
Veröffentlicht: (2002)
Veröffentlicht: (2002)
What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
von: Errica, Federico, et al.
Veröffentlicht: (2024)
von: Errica, Federico, et al.
Veröffentlicht: (2024)
Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
Towards Understanding Sycophancy in Language Models
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
von: Sharma, Mrinank, et al.
Veröffentlicht: (2023)
Extending Beacon to Hindi: Cultural Adaptation Drives Cross-Lingual Sycophancy
von: Sattigeri, Sarthak
Veröffentlicht: (2026)
von: Sattigeri, Sarthak
Veröffentlicht: (2026)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
von: Çelebi, Yusuf, et al.
Veröffentlicht: (2025)
von: Çelebi, Yusuf, et al.
Veröffentlicht: (2025)
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2026)
When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
von: Zolfaghari, Vahideh
Veröffentlicht: (2026)
Out-of-Distribution Detection Methods Answer the Wrong Questions
von: Li, Yucen Lily, et al.
Veröffentlicht: (2025)
von: Li, Yucen Lily, et al.
Veröffentlicht: (2025)
Your Assumed DAG is Wrong and Here's How To Deal With It
von: Padh, Kirtan, et al.
Veröffentlicht: (2025)
von: Padh, Kirtan, et al.
Veröffentlicht: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
von: Shi, Claudia, et al.
Veröffentlicht: (2024)
Capacity-Aware Planning and Scheduling in Budget-Constrained Multi-Agent MDPs: A Meta-RL Approach
von: Vora, Manav, et al.
Veröffentlicht: (2024)
von: Vora, Manav, et al.
Veröffentlicht: (2024)
To Agree or To Be Right? The Grounding-Sycophancy Tradeoff in Medical Vision-Language Models
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
CoDec: Prefix-Shared Decoding Kernel for LLMs
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
von: Wang, Zhibin, et al.
Veröffentlicht: (2025)
Physicists Don't Know What They're Talking About When They Say 'Order'
von: Arafat Gaspar Jiménez Gaistardo
Veröffentlicht: (2025)
von: Arafat Gaspar Jiménez Gaistardo
Veröffentlicht: (2025)
Thinking Out Loud: Do Reasoning Models Know When They're Right?
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
von: Zeng, Qingcheng, et al.
Veröffentlicht: (2025)
Data Unlearning in Diffusion Models
von: Alberti, Silas, et al.
Veröffentlicht: (2025)
von: Alberti, Silas, et al.
Veröffentlicht: (2025)
KnowCoder: Coding Structured Knowledge into LLMs for Universal Information Extraction
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
von: Li, Zixuan, et al.
Veröffentlicht: (2024)
They're happy and they know it
von: Anscombe, Nadya
Veröffentlicht: (2005)
von: Anscombe, Nadya
Veröffentlicht: (2005)
They're redesigning the airplane
Veröffentlicht: (1981)
Veröffentlicht: (1981)
Ähnliche Einträge
-
Do Androids Know They're Only Dreaming of Electric Sheep?
von: CH-Wang, Sky, et al.
Veröffentlicht: (2023) -
Know When You're Wrong: Aligning Confidence with Correctness for LLM Error Detection
von: Xiaohu, Xie, et al.
Veröffentlicht: (2026) -
AgreeMate: Teaching LLMs to Haggle
von: Chatterjee, Ainesh, et al.
Veröffentlicht: (2024) -
Letting Others Know How They're Doing.
von: Hartzell, Gary
Veröffentlicht: (1993) -
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
von: Magid, Salma Abdel, et al.
Veröffentlicht: (2024)