Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Shah, Arya, Mishra, Deepali, Silpasuwanchai, Chaklam |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters
por: Shah, Arya, et al.
Publicado: (2026)
por: Shah, Arya, et al.
Publicado: (2026)
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
por: Shah, Arya, et al.
Publicado: (2026)
por: Shah, Arya, et al.
Publicado: (2026)
From Overload to Convergence: Supporting Multi-Issue Human-AI Negotiation with Bayesian Visualization
por: Parmar, Mehul, et al.
Publicado: (2026)
por: Parmar, Mehul, et al.
Publicado: (2026)
Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights
por: Mishra, Deepali, et al.
Publicado: (2025)
por: Mishra, Deepali, et al.
Publicado: (2025)
To Tell The Truth: Language of Deception and Language Models
por: Hazra, Sanchaita, et al.
Publicado: (2023)
por: Hazra, Sanchaita, et al.
Publicado: (2023)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
por: Barkett, Emilio, et al.
Publicado: (2025)
por: Barkett, Emilio, et al.
Publicado: (2025)
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains
por: Yi, Zihao, et al.
Publicado: (2025)
por: Yi, Zihao, et al.
Publicado: (2025)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
por: Natan, Shahar Ben, et al.
Publicado: (2026)
por: Natan, Shahar Ben, et al.
Publicado: (2026)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
por: Kasprova, Vira, et al.
Publicado: (2026)
por: Kasprova, Vira, et al.
Publicado: (2026)
Sycophancy in Large Language Models: Causes and Mitigations
por: Malmqvist, Lars
Publicado: (2024)
por: Malmqvist, Lars
Publicado: (2024)
Measuring Agreeableness Bias in Multimodal Models
por: Lim, Jaehyuk, et al.
Publicado: (2024)
por: Lim, Jaehyuk, et al.
Publicado: (2024)
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay
por: de Carvalho, Gonçalo Hora, et al.
Publicado: (2024)
por: de Carvalho, Gonçalo Hora, et al.
Publicado: (2024)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
por: Denison, Carson, et al.
Publicado: (2024)
por: Denison, Carson, et al.
Publicado: (2024)
Quantifying Risk Propensities of Large Language Models: Ethical Focus and Bias Detection through Role-Play
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Accounting for Sycophancy in Language Model Uncertainty Estimation
por: Sicilia, Anthony, et al.
Publicado: (2024)
por: Sicilia, Anthony, et al.
Publicado: (2024)
Role-Playing Evaluation for Large Language Models
por: Boudouri, Yassine El, et al.
Publicado: (2025)
por: Boudouri, Yassine El, et al.
Publicado: (2025)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
por: Pandey, Sanskar, et al.
Publicado: (2025)
por: Pandey, Sanskar, et al.
Publicado: (2025)
Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework
por: Zhao, Yunpu, et al.
Publicado: (2024)
por: Zhao, Yunpu, et al.
Publicado: (2024)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
por: Shah, Arya, et al.
Publicado: (2026)
por: Shah, Arya, et al.
Publicado: (2026)
Towards Understanding Sycophancy in Language Models
por: Sharma, Mrinank, et al.
Publicado: (2023)
por: Sharma, Mrinank, et al.
Publicado: (2023)
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models
por: Wang, Zekun Moore, et al.
Publicado: (2023)
por: Wang, Zekun Moore, et al.
Publicado: (2023)
Sycophancy Hides Linearly in the Attention Heads
por: Genadi, Rifo, et al.
Publicado: (2026)
por: Genadi, Rifo, et al.
Publicado: (2026)
BASIL: Bayesian Assessment of Sycophancy in LLMs
por: Atwell, Katherine, et al.
Publicado: (2025)
por: Atwell, Katherine, et al.
Publicado: (2025)
Role-Playing Agents Driven by Large Language Models: Current Status, Challenges, and Future Trends
por: Wang, Ye, et al.
Publicado: (2026)
por: Wang, Ye, et al.
Publicado: (2026)
The Oscars of AI Theater: A Survey on Role-Playing with Language Models
por: Chen, Nuo, et al.
Publicado: (2024)
por: Chen, Nuo, et al.
Publicado: (2024)
RPGBENCH: Evaluating Large Language Models as Role-Playing Game Engines
por: Yu, Pengfei, et al.
Publicado: (2025)
por: Yu, Pengfei, et al.
Publicado: (2025)
On the Decision-Making Abilities in Role-Playing using Large Language Models
por: Shen, Chenglei, et al.
Publicado: (2024)
por: Shen, Chenglei, et al.
Publicado: (2024)
Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
por: Christian, Brian, et al.
Publicado: (2026)
por: Christian, Brian, et al.
Publicado: (2026)
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing
por: Xu, Jiacheng, et al.
Publicado: (2026)
por: Xu, Jiacheng, et al.
Publicado: (2026)
PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
por: Çelebi, Yusuf, et al.
Publicado: (2025)
por: Çelebi, Yusuf, et al.
Publicado: (2025)
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs
por: Wang, Kai, et al.
Publicado: (2026)
por: Wang, Kai, et al.
Publicado: (2026)
Orca: Enhancing Role-Playing Abilities of Large Language Models by Integrating Personality Traits
por: Huang, Yuxuan
Publicado: (2024)
por: Huang, Yuxuan
Publicado: (2024)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
por: Chen, Zhongzhi, et al.
Publicado: (2023)
por: Chen, Zhongzhi, et al.
Publicado: (2023)
On the Relationship between Truth and Political Bias in Language Models
por: Fulay, Suyash, et al.
Publicado: (2024)
por: Fulay, Suyash, et al.
Publicado: (2024)
RoleCraft-GLM: Advancing Personalized Role-Playing in Large Language Models
por: Tao, Meiling, et al.
Publicado: (2023)
por: Tao, Meiling, et al.
Publicado: (2023)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
por: Costa, Davi Bastos, et al.
Publicado: (2025)
por: Costa, Davi Bastos, et al.
Publicado: (2025)
Truth Knows No Language: Evaluating Truthfulness Beyond English
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
por: Figueras, Blanca Calvo, et al.
Publicado: (2025)
Tell Me More! Towards Implicit User Intention Understanding of Language Model Driven Agents
por: Qian, Cheng, et al.
Publicado: (2024)
por: Qian, Cheng, et al.
Publicado: (2024)
Rehearse With User: Personalized Opinion Summarization via Role-Playing based on Large Language Models
por: Zhang, Yanyue, et al.
Publicado: (2025)
por: Zhang, Yanyue, et al.
Publicado: (2025)
LLM Discussion: Enhancing the Creativity of Large Language Models via Discussion Framework and Role-Play
por: Lu, Li-Chun, et al.
Publicado: (2024)
por: Lu, Li-Chun, et al.
Publicado: (2024)
Ejemplares similares
-
SycoPhantasy: Quantifying Sycophancy and Hallucination in Small Open Weight VLMs for Vision-Language Scoring of Fantasy Characters
por: Shah, Arya, et al.
Publicado: (2026) -
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
por: Shah, Arya, et al.
Publicado: (2026) -
From Overload to Convergence: Supporting Multi-Issue Human-AI Negotiation with Bayesian Visualization
por: Parmar, Mehul, et al.
Publicado: (2026) -
Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights
por: Mishra, Deepali, et al.
Publicado: (2025) -
To Tell The Truth: Language of Deception and Language Models
por: Hazra, Sanchaita, et al.
Publicado: (2023)