Saved in:
Bibliographic Details
Main Authors: Villa-Cueva, Emilio, Ahmed, S M Masrur, Chevi, Rendi, Cruz, Jan Christian Blaise, Elzeky, Kareem, Cristobal, Fermin, Aji, Alham Fikri, Wang, Skyler, Mihalcea, Rada, Solorio, Thamar
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2507.04415
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916959623315456
author Villa-Cueva, Emilio
Ahmed, S M Masrur
Chevi, Rendi
Cruz, Jan Christian Blaise
Elzeky, Kareem
Cristobal, Fermin
Aji, Alham Fikri
Wang, Skyler
Mihalcea, Rada
Solorio, Thamar
author_facet Villa-Cueva, Emilio
Ahmed, S M Masrur
Chevi, Rendi
Cruz, Jan Christian Blaise
Elzeky, Kareem
Cristobal, Fermin
Aji, Alham Fikri
Wang, Skyler
Mihalcea, Rada
Solorio, Thamar
contents Understanding Theory of Mind is essential for building socially intelligent multimodal agents capable of perceiving and interpreting human behavior. We introduce MoMentS (Multimodal Mental States), a comprehensive benchmark designed to assess the ToM capabilities of multimodal large language models (LLMs) through realistic, narrative-rich scenarios presented in short films. MoMentS includes over 2,300 multiple-choice questions spanning seven distinct ToM categories. The benchmark features long video context windows and realistic social interactions that provide deeper insight into characters' mental states. We evaluate several MLLMs and find that although vision generally improves performance, models still struggle to integrate it effectively. For audio, models that process dialogues as audio do not consistently outperform transcript-based inputs. Our findings highlight the need to improve multimodal integration and point to open challenges that must be addressed to advance AI's social understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04415
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
Villa-Cueva, Emilio
Ahmed, S M Masrur
Chevi, Rendi
Cruz, Jan Christian Blaise
Elzeky, Kareem
Cristobal, Fermin
Aji, Alham Fikri
Wang, Skyler
Mihalcea, Rada
Solorio, Thamar
Computation and Language
Understanding Theory of Mind is essential for building socially intelligent multimodal agents capable of perceiving and interpreting human behavior. We introduce MoMentS (Multimodal Mental States), a comprehensive benchmark designed to assess the ToM capabilities of multimodal large language models (LLMs) through realistic, narrative-rich scenarios presented in short films. MoMentS includes over 2,300 multiple-choice questions spanning seven distinct ToM categories. The benchmark features long video context windows and realistic social interactions that provide deeper insight into characters' mental states. We evaluate several MLLMs and find that although vision generally improves performance, models still struggle to integrate it effectively. For audio, models that process dialogues as audio do not consistently outperform transcript-based inputs. Our findings highlight the need to improve multimodal integration and point to open challenges that must be addressed to advance AI's social understanding.
title MOMENTS: A Comprehensive Multimodal Benchmark for Theory of Mind
topic Computation and Language
url https://arxiv.org/abs/2507.04415