Training language models to be warm and empathetic makes them less reliable and more sycophantic
Fuente:
arXiv
Saved in:
| Main Authors: | Ibrahim, Lujain, Hafner, Franziska Sofia, Rocher, Luc |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
by: Ibrahim, Lujain, et al.
Published: (2026)
by: Ibrahim, Lujain, et al.
Published: (2026)
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
by: Hafner, Franziska Sofia, et al.
Published: (2025)
by: Hafner, Franziska Sofia, et al.
Published: (2025)
Characterizing and modeling harms from interactions with design patterns in AI interfaces
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
Developmental trajectories of decision making and affective dynamics in large language models
by: Wang, Zhihao, et al.
Published: (2025)
by: Wang, Zhihao, et al.
Published: (2025)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
Verbalizing LLMs' assumptions to explain and control sycophancy
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Uncovering inequalities in new knowledge learning by large language models across different languages
by: Wang, Chenglong, et al.
Published: (2025)
by: Wang, Chenglong, et al.
Published: (2025)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
Do Chinese models speak Chinese languages?
by: Wen-Yi, Andrea W, et al.
Published: (2025)
by: Wen-Yi, Andrea W, et al.
Published: (2025)
Failure of contextual invariance in large language models
by: Kumar, Sagar, et al.
Published: (2026)
by: Kumar, Sagar, et al.
Published: (2026)
Large language models in medicine: the potentials and pitfalls
by: Omiye, Jesutofunmi A., et al.
Published: (2023)
by: Omiye, Jesutofunmi A., et al.
Published: (2023)
What can large language models do for sustainable food?
by: Thomas, Anna T., et al.
Published: (2025)
by: Thomas, Anna T., et al.
Published: (2025)
A closer look at how large language models trust humans: patterns and biases
by: Lerman, Valeria, et al.
Published: (2025)
by: Lerman, Valeria, et al.
Published: (2025)
A survey on fairness of large language models in e-commerce: progress, application, and challenge
by: Ren, Qingyang, et al.
Published: (2024)
by: Ren, Qingyang, et al.
Published: (2024)
Large language models can consistently generate high-quality content for election disinformation operations
by: Williams, Angus R., et al.
Published: (2024)
by: Williams, Angus R., et al.
Published: (2024)
AI-AI Bias: large language models favor communications generated by large language models
by: Laurito, Walter, et al.
Published: (2024)
by: Laurito, Walter, et al.
Published: (2024)
Redefining technology for indigenous languages
by: Fernandez-Sabido, Silvia, et al.
Published: (2025)
by: Fernandez-Sabido, Silvia, et al.
Published: (2025)
Meaning Is Not A Metric: Using LLMs to make cultural context legible at scale
by: Kommers, Cody, et al.
Published: (2025)
by: Kommers, Cody, et al.
Published: (2025)
Can LLMs make trade-offs involving stipulated pain and pleasure states?
by: Keeling, Geoff, et al.
Published: (2024)
by: Keeling, Geoff, et al.
Published: (2024)
Can adversarial attacks by large language models be attributed?
by: Cebrian, Manuel, et al.
Published: (2024)
by: Cebrian, Manuel, et al.
Published: (2024)
HumT DumT: Measuring and controlling human-like language in LLMs
by: Cheng, Myra, et al.
Published: (2025)
by: Cheng, Myra, et al.
Published: (2025)
Developing Story: Case Studies of Generative AI's Use in Journalism
by: Brigham, Natalie Grace, et al.
Published: (2024)
by: Brigham, Natalie Grace, et al.
Published: (2024)
Implicit assessment of language learning during practice as accurate as explicit testing
by: Hou, Jue, et al.
Published: (2024)
by: Hou, Jue, et al.
Published: (2024)
Profiling learners' affective engagement: Emotion AI, intercultural pragmatics, and language learning
by: Godwin-Jones, Robert
Published: (2026)
by: Godwin-Jones, Robert
Published: (2026)
Assessing the nature of large language models: A caution against anthropocentrism
by: Speed, Ann
Published: (2023)
by: Speed, Ann
Published: (2023)
A validity-guided workflow for robust large language model research in psychology
by: Lin, Zhicheng
Published: (2025)
by: Lin, Zhicheng
Published: (2025)
Evidence of a log scaling law for political persuasion with large language models
by: Hackenburg, Kobi, et al.
Published: (2024)
by: Hackenburg, Kobi, et al.
Published: (2024)
Towards interactive evaluations for interaction harms in human-AI systems
by: Ibrahim, Lujain, et al.
Published: (2024)
by: Ibrahim, Lujain, et al.
Published: (2024)
What are human values, and how do we align AI to them?
by: Klingefjord, Oliver, et al.
Published: (2024)
by: Klingefjord, Oliver, et al.
Published: (2024)
Human Decision-making is Susceptible to AI-driven Manipulation
by: Sabour, Sahand, et al.
Published: (2025)
by: Sabour, Sahand, et al.
Published: (2025)
WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis
by: Wu, Yuqi, et al.
Published: (2025)
by: Wu, Yuqi, et al.
Published: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
The Last Fingerprint: How Markdown Training Shapes LLM Prose
by: Freeburg, E. M.
Published: (2026)
by: Freeburg, E. M.
Published: (2026)
PersLLM: A Personified Training Approach for Large Language Models
by: Zeng, Zheni, et al.
Published: (2024)
by: Zeng, Zheni, et al.
Published: (2024)
From Hard Refusals to Safe-Completions: Toward Output-Centric Safety Training
by: Yuan, Yuan, et al.
Published: (2025)
by: Yuan, Yuan, et al.
Published: (2025)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
by: Kiet, Huynh Trung, et al.
Published: (2026)
by: Kiet, Huynh Trung, et al.
Published: (2026)
The opportunities and risks of large language models in mental health
by: Lawrence, Hannah R., et al.
Published: (2024)
by: Lawrence, Hannah R., et al.
Published: (2024)
Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers
by: Chakrabarty, Tuhin, et al.
Published: (2025)
by: Chakrabarty, Tuhin, et al.
Published: (2025)
Can a large language model be a gaslighter?
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Similar Items
-
Sycophantic AI makes human interaction feel more effortful and less satisfying over time
by: Ibrahim, Lujain, et al.
Published: (2026) -
Gender Trouble in Language Models: An Empirical Audit Guided by Gender Performativity Theory
by: Hafner, Franziska Sofia, et al.
Published: (2025) -
Characterizing and modeling harms from interactions with design patterns in AI interfaces
by: Ibrahim, Lujain, et al.
Published: (2024) -
Developmental trajectories of decision making and affective dynamics in large language models
by: Wang, Zhihao, et al.
Published: (2025) -
ELEPHANT: Measuring and understanding social sycophancy in LLMs
by: Cheng, Myra, et al.
Published: (2025)