Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Zhawnen, Wang, Tianchun, Wang, Yizhou, Kosinski, Michal, Zhang, Xiang, Fu, Yun, Li, Sheng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916947989364736
author Chen, Zhawnen
Wang, Tianchun
Wang, Yizhou
Kosinski, Michal
Zhang, Xiang
Fu, Yun
Li, Sheng
author_facet Chen, Zhawnen
Wang, Tianchun
Wang, Yizhou
Kosinski, Michal
Zhang, Xiang
Fu, Yun
Li, Sheng
contents Can large multimodal models have a human-like ability for emotional and social reasoning, and if so, how does it work? Recent research has discovered emergent theory-of-mind (ToM) reasoning capabilities in large language models (LLMs). LLMs can reason about people's mental states by solving various text-based ToM tasks that ask questions about the actors' ToM (e.g., human belief, desire, intention). However, human reasoning in the wild is often grounded in dynamic scenes across time. Thus, we consider videos a new medium for examining spatio-temporal ToM reasoning ability. Specifically, we ask explicit probing questions about videos with abundant social and emotional reasoning content. We develop a pipeline for multimodal LLM for ToM reasoning using video and text. We also enable explicit ToM reasoning by retrieving key frames for answering a ToM question, which reveals how multimodal LLMs reason about ToM.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13763
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
Chen, Zhawnen
Wang, Tianchun
Wang, Yizhou
Kosinski, Michal
Zhang, Xiang
Fu, Yun
Li, Sheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Can large multimodal models have a human-like ability for emotional and social reasoning, and if so, how does it work? Recent research has discovered emergent theory-of-mind (ToM) reasoning capabilities in large language models (LLMs). LLMs can reason about people's mental states by solving various text-based ToM tasks that ask questions about the actors' ToM (e.g., human belief, desire, intention). However, human reasoning in the wild is often grounded in dynamic scenes across time. Thus, we consider videos a new medium for examining spatio-temporal ToM reasoning ability. Specifically, we ask explicit probing questions about videos with abundant social and emotional reasoning content. We develop a pipeline for multimodal LLM for ToM reasoning using video and text. We also enable explicit ToM reasoning by retrieving key frames for answering a ToM question, which reveals how multimodal LLMs reason about ToM.
title Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2406.13763