Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Lubrano, Kate M., Sayed, Faisal, Rathod, Ankita, Akshansh, Thomas-Smith, Craver Corbyn, Whiting, Mark E., Nguyen, Karina
Format:	Preprint
Published:	2026
Subjects:	Artificial Intelligence
Online Access:	https://arxiv.org/abs/2605.21739
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866911725963444224
author	Lubrano, Kate M. Sayed, Faisal Rathod, Ankita Akshansh Thomas-Smith, Craver Corbyn Whiting, Mark E. Nguyen, Karina
author_facet	Lubrano, Kate M. Sayed, Faisal Rathod, Ankita Akshansh Thomas-Smith, Craver Corbyn Whiting, Mark E. Nguyen, Karina
contents	Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly important to assess as LLMs assume conversational roles in everyday life. Existing EI benchmarks rely on synthetic prompts, single-turn cases, or third-party annotation. These approaches do not directly measure how models infer and respond to a participant's emotional state over the course of a real conversation. We introduce AttuneBench, a benchmark grounded in 200 genuine multi-turn human-model conversations in which participants conversed with anonymized LLMs and provided turn-by-turn annotations of their emotional state, the model's behavior, and their preferred responses. Across 11 evaluated models, we find that model rankings on emotion recognition, behavioral classification, preference prediction, and judged response quality are largely independent, indicating that emotionally intelligent behavior decomposes into separable capabilities. Preference alignment and response-quality judgments are substantially more model-discriminating than emotion-label accuracy. These results indicate that emotionally intelligent behavior requires predicting what kind of response a specific user wants in context, a distinction that aggregate scoring can obscure and that single-turn or synthetic formats cannot directly capture across turns. AttuneBench provides a framework for assessing each of these capabilities and for diagnosing model-specific strengths and failure modes in emotionally salient conversation.
format	Preprint
id	arxiv_https___arxiv_org_abs_2605_21739
institution	arXiv
publishDate	2026
record_format	arxiv
spellingShingle	AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence Lubrano, Kate M. Sayed, Faisal Rathod, Ankita Akshansh Thomas-Smith, Craver Corbyn Whiting, Mark E. Nguyen, Karina Artificial Intelligence Emotional intelligence (EI), the ability to perceive, understand, and respond appropriately to others' emotional states, is central to human communication, and increasingly important to assess as LLMs assume conversational roles in everyday life. Existing EI benchmarks rely on synthetic prompts, single-turn cases, or third-party annotation. These approaches do not directly measure how models infer and respond to a participant's emotional state over the course of a real conversation. We introduce AttuneBench, a benchmark grounded in 200 genuine multi-turn human-model conversations in which participants conversed with anonymized LLMs and provided turn-by-turn annotations of their emotional state, the model's behavior, and their preferred responses. Across 11 evaluated models, we find that model rankings on emotion recognition, behavioral classification, preference prediction, and judged response quality are largely independent, indicating that emotionally intelligent behavior decomposes into separable capabilities. Preference alignment and response-quality judgments are substantially more model-discriminating than emotion-label accuracy. These results indicate that emotionally intelligent behavior requires predicting what kind of response a specific user wants in context, a distinction that aggregate scoring can obscure and that single-turn or synthetic formats cannot directly capture across turns. AttuneBench provides a framework for assessing each of these capabilities and for diagnosing model-specific strengths and failure modes in emotionally salient conversation.
title	AttuneBench: A Conversation-Based Benchmark for LLM Emotional Intelligence
topic	Artificial Intelligence
url	https://arxiv.org/abs/2605.21739

Similar Items