EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Paech, Samuel J.
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917557966995456
author Paech, Samuel J.
author_facet Paech, Samuel J.
contents We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com
format Preprint
id arxiv_https___arxiv_org_abs_2312_06281
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Paech, Samuel J.
Computation and Language
Artificial Intelligence
I.2.7
We introduce EQ-Bench, a novel benchmark designed to evaluate aspects of emotional intelligence in Large Language Models (LLMs). We assess the ability of LLMs to understand complex emotions and social interactions by asking them to predict the intensity of emotional states of characters in a dialogue. The benchmark is able to discriminate effectively between a wide range of models. We find that EQ-Bench correlates strongly with comprehensive multi-domain benchmarks like MMLU (Hendrycks et al., 2020) (r=0.97), indicating that we may be capturing similar aspects of broad intelligence. Our benchmark produces highly repeatable results using a set of 60 English-language questions. We also provide open-source code for an automated benchmarking pipeline at https://github.com/EQ-bench/EQ-Bench and a leaderboard at https://eqbench.com
title EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
topic Computation and Language
Artificial Intelligence
I.2.7
url https://arxiv.org/abs/2312.06281