Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Asanuma, Haruka, Koide-Majima, Naoko, Nakamura, Ken, Horii, Takato, Nishimoto, Shinji, Oizumi, Masafumi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913854446895104
author Asanuma, Haruka
Koide-Majima, Naoko
Nakamura, Ken
Horii, Takato
Nishimoto, Shinji
Oizumi, Masafumi
author_facet Asanuma, Haruka
Koide-Majima, Naoko
Nakamura, Ken
Horii, Takato
Nishimoto, Shinji
Oizumi, Masafumi
contents Recent studies have revealed that human emotions exhibit a high-dimensional, complex structure. A full capturing of this complexity requires new approaches, as conventional models that disregard high dimensionality risk overlooking key nuances of human emotions. Here, we examined the extent to which the latest generation of rapidly evolving Multimodal Large Language Models (MLLMs) capture these high-dimensional, intricate emotion structures, including capabilities and limitations. Specifically, we compared self-reported emotion ratings from participants watching videos with model-generated estimates (e.g., Gemini or GPT). We evaluated performance not only at the individual video level but also from emotion structures that account for inter-video relationships. At the level of simple correlation between emotion structures, our results demonstrated strong similarity between human and model-inferred emotion structures. To further explore whether the similarity between humans and models is at the signle item level or the coarse-categorical level, we applied Gromov Wasserstein Optimal Transport. We found that although performance was not necessarily high at the strict, single-item level, performance across video categories that elicit similar emotions was substantial, indicating that the model could infer human emotional experiences at the category level. Our results suggest that current state-of-the-art MLLMs broadly capture the complex high-dimensional emotion structures at the category level, as well as their apparent limitations in accurately capturing entire structures at the single-item level.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12746
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
Asanuma, Haruka
Koide-Majima, Naoko
Nakamura, Ken
Horii, Takato
Nishimoto, Shinji
Oizumi, Masafumi
Artificial Intelligence
I.2.7; I.2.10; I.5.1
Recent studies have revealed that human emotions exhibit a high-dimensional, complex structure. A full capturing of this complexity requires new approaches, as conventional models that disregard high dimensionality risk overlooking key nuances of human emotions. Here, we examined the extent to which the latest generation of rapidly evolving Multimodal Large Language Models (MLLMs) capture these high-dimensional, intricate emotion structures, including capabilities and limitations. Specifically, we compared self-reported emotion ratings from participants watching videos with model-generated estimates (e.g., Gemini or GPT). We evaluated performance not only at the individual video level but also from emotion structures that account for inter-video relationships. At the level of simple correlation between emotion structures, our results demonstrated strong similarity between human and model-inferred emotion structures. To further explore whether the similarity between humans and models is at the signle item level or the coarse-categorical level, we applied Gromov Wasserstein Optimal Transport. We found that although performance was not necessarily high at the strict, single-item level, performance across video categories that elicit similar emotions was substantial, indicating that the model could infer human emotional experiences at the category level. Our results suggest that current state-of-the-art MLLMs broadly capture the complex high-dimensional emotion structures at the category level, as well as their apparent limitations in accurately capturing entire structures at the single-item level.
title Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
topic Artificial Intelligence
I.2.7; I.2.10; I.5.1
url https://arxiv.org/abs/2505.12746