When Large Language Models are Reliable for Judging Empathic Communication

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kumar, Aakriti, Poungpeth, Nalin, Yang, Diyi, Farrell, Erina, Lambert, Bruce, Groh, Matthew
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911189945024512
author Kumar, Aakriti
Poungpeth, Nalin
Yang, Diyi
Farrell, Erina
Lambert, Bruce
Groh, Matthew
author_facet Kumar, Aakriti
Poungpeth, Nalin
Yang, Diyi
Farrell, Erina
Lambert, Bruce
Groh, Matthew
contents Large language models (LLMs) excel at generating empathic responses in text-based conversations. But, how reliably do they judge the nuances of empathic communication? We investigate this question by comparing how experts, crowdworkers, and LLMs annotate empathic communication across four evaluative frameworks drawn from psychology, natural language processing, and communications applied to 200 real-world conversations where one speaker shares a personal problem and the other offers support. Drawing on 3,150 expert annotations, 2,844 crowd annotations, and 3,150 LLM annotations, we assess inter-rater reliability between these three annotator groups. We find that expert agreement is high but varies across the frameworks' sub-components depending on their clarity, complexity, and subjectivity. We show that expert agreement offers a more informative benchmark for contextualizing LLM performance than standard classification metrics. Across all four frameworks, LLMs consistently approach this expert level benchmark and exceed the reliability of crowdworkers. These results demonstrate how LLMs, when validated on specific tasks with appropriate benchmarks, can support transparency and oversight in emotionally sensitive applications including their use as conversational companions.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10150
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Large Language Models are Reliable for Judging Empathic Communication
Kumar, Aakriti
Poungpeth, Nalin
Yang, Diyi
Farrell, Erina
Lambert, Bruce
Groh, Matthew
Computation and Language
Human-Computer Interaction
Large language models (LLMs) excel at generating empathic responses in text-based conversations. But, how reliably do they judge the nuances of empathic communication? We investigate this question by comparing how experts, crowdworkers, and LLMs annotate empathic communication across four evaluative frameworks drawn from psychology, natural language processing, and communications applied to 200 real-world conversations where one speaker shares a personal problem and the other offers support. Drawing on 3,150 expert annotations, 2,844 crowd annotations, and 3,150 LLM annotations, we assess inter-rater reliability between these three annotator groups. We find that expert agreement is high but varies across the frameworks' sub-components depending on their clarity, complexity, and subjectivity. We show that expert agreement offers a more informative benchmark for contextualizing LLM performance than standard classification metrics. Across all four frameworks, LLMs consistently approach this expert level benchmark and exceed the reliability of crowdworkers. These results demonstrate how LLMs, when validated on specific tasks with appropriate benchmarks, can support transparency and oversight in emotionally sensitive applications including their use as conversational companions.
title When Large Language Models are Reliable for Judging Empathic Communication
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2506.10150