Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Sungwon, Khashabi, Daniel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Evaluating the Evaluators: Are readability metrics good measures of readability?
di: Cachola, Isabel, et al.
Pubblicazione: (2025)
di: Cachola, Isabel, et al.
Pubblicazione: (2025)
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
di: He, Zhitao, et al.
Pubblicazione: (2026)
di: He, Zhitao, et al.
Pubblicazione: (2026)
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
di: Byerly, Adam, et al.
Pubblicazione: (2025)
di: Byerly, Adam, et al.
Pubblicazione: (2025)
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
di: Li, Tianjian, et al.
Pubblicazione: (2025)
di: Li, Tianjian, et al.
Pubblicazione: (2025)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
di: Byerly, Adam, et al.
Pubblicazione: (2024)
di: Byerly, Adam, et al.
Pubblicazione: (2024)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
di: Wang, Weiqi, et al.
Pubblicazione: (2025)
di: Wang, Weiqi, et al.
Pubblicazione: (2025)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
di: Nogueira, Rodrigo, et al.
Pubblicazione: (2026)
di: Nogueira, Rodrigo, et al.
Pubblicazione: (2026)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
di: Hong, Jiseung, et al.
Pubblicazione: (2025)
di: Hong, Jiseung, et al.
Pubblicazione: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
di: Natan, Shahar Ben, et al.
Pubblicazione: (2026)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
di: Lu, Taiming, et al.
Pubblicazione: (2024)
di: Lu, Taiming, et al.
Pubblicazione: (2024)
Certified Mitigation of Worst-Case LLM Copyright Infringement
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
di: Zhang, Jingyu, et al.
Pubblicazione: (2025)
Hell or High Water: Evaluating Agentic Recovery from External Failures
di: Wang, Andrew, et al.
Pubblicazione: (2025)
di: Wang, Andrew, et al.
Pubblicazione: (2025)
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2025)
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2025)
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
di: Feng, Zhaoxin, et al.
Pubblicazione: (2026)
di: Feng, Zhaoxin, et al.
Pubblicazione: (2026)
RbtAct: Rebuttal as Supervision for Actionable Review Feedback Generation
di: Wu, Sihong, et al.
Pubblicazione: (2026)
di: Wu, Sihong, et al.
Pubblicazione: (2026)
RORA: Robust Free-Text Rationale Evaluation
di: Jiang, Zhengping, et al.
Pubblicazione: (2024)
di: Jiang, Zhengping, et al.
Pubblicazione: (2024)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
di: Uzunoglu, Arda, et al.
Pubblicazione: (2026)
di: Uzunoglu, Arda, et al.
Pubblicazione: (2026)
It's Not Always Sycophancy: Measuring LLM Conformity as a Function of Epistemic Uncertainty
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
di: Guo, Kevin H., et al.
Pubblicazione: (2026)
Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
BASIL: Bayesian Assessment of Sycophancy in LLMs
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
di: Atwell, Katherine, et al.
Pubblicazione: (2025)
Sycophancy Hides Linearly in the Attention Heads
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
di: Genadi, Rifo, et al.
Pubblicazione: (2026)
Many-Tier Instruction Hierarchy in LLM Agents
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
di: Zhang, Jingyu, et al.
Pubblicazione: (2026)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
Peacemaker or Troublemaker: How Sycophancy Shapes Multi-Agent Debate
di: Yao, Binwei, et al.
Pubblicazione: (2025)
di: Yao, Binwei, et al.
Pubblicazione: (2025)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
di: Liu, Joshua, et al.
Pubblicazione: (2025)
di: Liu, Joshua, et al.
Pubblicazione: (2025)
Sycophancy in Large Language Models: Causes and Mitigations
di: Malmqvist, Lars
Pubblicazione: (2024)
di: Malmqvist, Lars
Pubblicazione: (2024)
Sycophancy Claims about Language Models: The Missing Human-in-the-Loop
di: Batzner, Jan, et al.
Pubblicazione: (2025)
di: Batzner, Jan, et al.
Pubblicazione: (2025)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
di: Mishra, Aayush, et al.
Pubblicazione: (2025)
di: Mishra, Aayush, et al.
Pubblicazione: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
di: Shen, Lingfeng, et al.
Pubblicazione: (2023)
di: Shen, Lingfeng, et al.
Pubblicazione: (2023)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
di: Ou, Jiefu, et al.
Pubblicazione: (2024)
di: Ou, Jiefu, et al.
Pubblicazione: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
di: Wang, Hexuan, et al.
Pubblicazione: (2026)
di: Wang, Hexuan, et al.
Pubblicazione: (2026)
Accounting for Sycophancy in Language Model Uncertainty Estimation
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases
di: Koretsky, Mathew J., et al.
Pubblicazione: (2025)
di: Koretsky, Mathew J., et al.
Pubblicazione: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
di: Sahoo, Subramanyam
Pubblicazione: (2026)
di: Sahoo, Subramanyam
Pubblicazione: (2026)
"Check My Work?": Measuring Sycophancy in a Simulated Educational Context
di: Arvin, Chuck
Pubblicazione: (2025)
di: Arvin, Chuck
Pubblicazione: (2025)
SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
di: Park, Sungwon, et al.
Pubblicazione: (2024)
di: Park, Sungwon, et al.
Pubblicazione: (2024)
Crystal: Characterizing Relative Impact of Scholarly Publications
di: Collison, Hannah, et al.
Pubblicazione: (2026)
di: Collison, Hannah, et al.
Pubblicazione: (2026)
Feedback Friction: LLMs Struggle to Fully Incorporate External Feedback
di: Jiang, Dongwei, et al.
Pubblicazione: (2025)
di: Jiang, Dongwei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Evaluating the Evaluators: Are readability metrics good measures of readability?
di: Cachola, Isabel, et al.
Pubblicazione: (2025) -
RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind
di: He, Zhitao, et al.
Pubblicazione: (2026) -
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
di: Byerly, Adam, et al.
Pubblicazione: (2025) -
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning
di: Li, Tianjian, et al.
Pubblicazione: (2025) -
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
di: Byerly, Adam, et al.
Pubblicazione: (2024)