Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jen-tse, Lam, Man Ho, Li, Eric John, Ren, Shujie, Wang, Wenxuan, Jiao, Wenxiang, Tu, Zhaopeng, Lyu, Michael R.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909336661393408
author Huang, Jen-tse
Lam, Man Ho
Li, Eric John
Ren, Shujie
Wang, Wenxuan
Jiao, Wenxiang
Tu, Zhaopeng
Lyu, Michael R.
author_facet Huang, Jen-tse
Lam, Man Ho
Li, Eric John
Ren, Shujie
Wang, Wenxuan
Jiao, Wenxiang
Tu, Zhaopeng
Lyu, Michael R.
contents Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes seven LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4, Mixtral-8x22B, and LLaMA-3.1. We find that, despite several misalignments, LLMs can generally respond appropriately to certain situations. Nevertheless, they fall short in alignment with the emotional behaviors of human beings and cannot establish connections between similar situations. Our collected dataset of situations, the human evaluation results, and the code of our testing framework, i.e., EmotionBench, are publicly available at https://github.com/CUHK-ARISE/EmotionBench.
format Preprint
id arxiv_https___arxiv_org_abs_2308_03656
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
Huang, Jen-tse
Lam, Man Ho
Li, Eric John
Ren, Shujie
Wang, Wenxuan
Jiao, Wenxiang
Tu, Zhaopeng
Lyu, Michael R.
Computation and Language
Evaluating Large Language Models' (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes seven LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4, Mixtral-8x22B, and LLaMA-3.1. We find that, despite several misalignments, LLMs can generally respond appropriately to certain situations. Nevertheless, they fall short in alignment with the emotional behaviors of human beings and cannot establish connections between similar situations. Our collected dataset of situations, the human evaluation results, and the code of our testing framework, i.e., EmotionBench, are publicly available at https://github.com/CUHK-ARISE/EmotionBench.
title Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench
topic Computation and Language
url https://arxiv.org/abs/2308.03656