Training language models to be warm and empathetic makes them less reliable and more sycophantic

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ibrahim, Lujain, Hafner, Franziska Sofia, Rocher, Luc
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913965667254272
author Ibrahim, Lujain
Hafner, Franziska Sofia
Rocher, Luc
author_facet Ibrahim, Lujain
Hafner, Franziska Sofia
Rocher, Luc
contents Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.
format Preprint
id arxiv_https___arxiv_org_abs_2507_21919
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Training language models to be warm and empathetic makes them less reliable and more sycophantic
Ibrahim, Lujain
Hafner, Franziska Sofia
Rocher, Luc
Computation and Language
Artificial Intelligence
Computers and Society
Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.
title Training language models to be warm and empathetic makes them less reliable and more sycophantic
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2507.21919