Sycophancy under Pressure: Evaluating and Mitigating Sycophantic Bias via Adversarial Dialogues in Scientific QA
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Kaiwei, Jia, Qi, Chen, Zijian, Sun, Wei, Zhu, Xiangyang, Li, Chunyi, Zhu, Dandan, Zhai, Guangtao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025)
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2025)
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2025)
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
di: Chen, Zijian, et al.
Pubblicazione: (2025)
di: Chen, Zijian, et al.
Pubblicazione: (2025)
User-centric Subjective Leaderboard by Customizable Reward Modeling
di: Jia, Qi, et al.
Pubblicazione: (2025)
di: Jia, Qi, et al.
Pubblicazione: (2025)
UniDial-EvalKit: A Unified Toolkit for Evaluating Multi-Faceted Conversational Abilities
di: Jia, Qi, et al.
Pubblicazione: (2026)
di: Jia, Qi, et al.
Pubblicazione: (2026)
Generalizable Video Quality Assessment via Weak-to-Strong Learning
di: Cao, Linhan, et al.
Pubblicazione: (2025)
di: Cao, Linhan, et al.
Pubblicazione: (2025)
VQAThinker: Exploring Generalizable and Explainable Video Quality Assessment via Reinforcement Learning
di: Cao, Linhan, et al.
Pubblicazione: (2025)
di: Cao, Linhan, et al.
Pubblicazione: (2025)
Q-Mirror: Unlocking the Multi-Modal Potential of Scientific Text-Only QA Pairs
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics
di: Zhang, Kaiwei, et al.
Pubblicazione: (2024)
di: Zhang, Kaiwei, et al.
Pubblicazione: (2024)
Mesh Mamba: A Unified State Space Model for Saliency Prediction in Non-Textured and Textured Meshes
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
di: Zhang, Kaiwei, et al.
Pubblicazione: (2025)
QualiRAG: Retrieval-Augmented Generation for Visual Quality Understanding
di: Cao, Linhan, et al.
Pubblicazione: (2026)
di: Cao, Linhan, et al.
Pubblicazione: (2026)
Human-Centric Evaluation for Foundation Models
di: Guo, Yijin, et al.
Pubblicazione: (2025)
di: Guo, Yijin, et al.
Pubblicazione: (2025)
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion
di: Nogueira, Rodrigo, et al.
Pubblicazione: (2026)
di: Nogueira, Rodrigo, et al.
Pubblicazione: (2026)
Measuring Sycophancy of Language Models in Multi-turn Dialogues
di: Hong, Jiseung, et al.
Pubblicazione: (2025)
di: Hong, Jiseung, et al.
Pubblicazione: (2025)
UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN Manipulation
di: Zhou, Hanzhang, et al.
Pubblicazione: (2024)
di: Zhou, Hanzhang, et al.
Pubblicazione: (2024)
PriceSeer: Evaluating Large Language Models in Real-Time Stock Prediction
di: Liang, Bohan, et al.
Pubblicazione: (2025)
di: Liang, Bohan, et al.
Pubblicazione: (2025)
Automated Safety Benchmarking: A Multi-agent Pipeline for LVLMs
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
Efficient Face Image Quality Assessment via Self-training and Knowledge Distillation
di: Sun, Wei, et al.
Pubblicazione: (2025)
di: Sun, Wei, et al.
Pubblicazione: (2025)
FVA-RAG: Falsification-Verification Alignment for Mitigating Sycophantic Hallucinations
di: Ravishankara, Mayank
Pubblicazione: (2025)
di: Ravishankara, Mayank
Pubblicazione: (2025)
Sycophancy in Large Language Models: Causes and Mitigations
di: Malmqvist, Lars
Pubblicazione: (2024)
di: Malmqvist, Lars
Pubblicazione: (2024)
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
di: Shen, Ye, et al.
Pubblicazione: (2026)
di: Shen, Ye, et al.
Pubblicazione: (2026)
Mitigating the Bias of Large Language Model Evaluation
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
di: Zhou, Hongli, et al.
Pubblicazione: (2024)
Breaking Bias, Building Bridges: Evaluation and Mitigation of Social Biases in LLMs via Contact Hypothesis
di: Raj, Chahat, et al.
Pubblicazione: (2024)
di: Raj, Chahat, et al.
Pubblicazione: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
di: Zhu, Ziyi, et al.
Pubblicazione: (2026)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
di: Barkett, Emilio, et al.
Pubblicazione: (2025)
KidVis: Do Multimodal Large Language Models Possess the Visual Perceptual Capabilities of a 6-Year-Old?
di: Wang, Xianfeng, et al.
Pubblicazione: (2026)
di: Wang, Xianfeng, et al.
Pubblicazione: (2026)
DebateQA: Evaluating Question Answering on Debatable Knowledge
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
SafeSci: Safety Evaluation of Large Language Models in Science Domains and Beyond
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
di: Zhu, Xiangyang, et al.
Pubblicazione: (2026)
The Ever-Evolving Science Exam
di: Wang, Junying, et al.
Pubblicazione: (2025)
di: Wang, Junying, et al.
Pubblicazione: (2025)
Bias in Large Language Models: Origin, Evaluation, and Mitigation
di: Guo, Yufei, et al.
Pubblicazione: (2024)
di: Guo, Yufei, et al.
Pubblicazione: (2024)
DetectiveQA: Evaluating Long-Context Reasoning on Detective Novels
di: Xu, Zhe, et al.
Pubblicazione: (2024)
di: Xu, Zhe, et al.
Pubblicazione: (2024)
SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
di: Bhalla, Joy, et al.
Pubblicazione: (2026)
Revealing User Familiarity Bias in Task-Oriented Dialogue via Interactive Evaluation
di: Kim, Takyoung, et al.
Pubblicazione: (2023)
di: Kim, Takyoung, et al.
Pubblicazione: (2023)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
di: Kim, Sungwon, et al.
Pubblicazione: (2025)
di: Kim, Sungwon, et al.
Pubblicazione: (2025)
Investigating the Influence of Language on Sycophantic Behavior of Multilingual LLMs
di: Aldahlawi, Bayan Abdullah, et al.
Pubblicazione: (2026)
di: Aldahlawi, Bayan Abdullah, et al.
Pubblicazione: (2026)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
di: Zhang, Chaowei, et al.
Pubblicazione: (2026)
di: Zhang, Chaowei, et al.
Pubblicazione: (2026)
A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
di: Shen, Ye, et al.
Pubblicazione: (2025)
di: Shen, Ye, et al.
Pubblicazione: (2025)
AdvSumm: Adversarial Training for Bias Mitigation in Text Summarization
di: Gupta, Mukur, et al.
Pubblicazione: (2025)
di: Gupta, Mukur, et al.
Pubblicazione: (2025)
Scientific QA System with Verifiable Answers
di: Ljajić, Adela, et al.
Pubblicazione: (2024)
di: Ljajić, Adela, et al.
Pubblicazione: (2024)
Documenti analoghi
-
SafetyFlow: An Agent-Flow System for Automated LLM Safety Benchmarking
di: Zhu, Xiangyang, et al.
Pubblicazione: (2025) -
Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
di: Vennemeyer, Daniel, et al.
Pubblicazione: (2025) -
One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework
di: Jia, Qi, et al.
Pubblicazione: (2025) -
Evaluating from Benign to Dynamic Adversarial: A Squid Game for Large Language Models
di: Chen, Zijian, et al.
Pubblicazione: (2025) -
User-centric Subjective Leaderboard by Customizable Reward Modeling
di: Jia, Qi, et al.
Pubblicazione: (2025)