MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zheng, Xiao, Nuoqian, Chai, Qi, Ye, Deheng, Wang, Hao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908536418598912
author Zhang, Zheng
Xiao, Nuoqian
Chai, Qi
Ye, Deheng
Wang, Hao
author_facet Zhang, Zheng
Xiao, Nuoqian
Chai, Qi
Ye, Deheng
Wang, Hao
contents Large Language Model (LLM) agents have demonstrated impressive capabilities in social deduction games (SDGs) like Werewolf, where strategic reasoning and social deception are essential. However, current approaches remain limited to textual information, ignoring crucial multimodal cues such as facial expressions and tone of voice that humans naturally use to communicate. Moreover, existing SDG agents primarily focus on inferring other players' identities without modeling how others perceive themselves or fellow players. To address these limitations, we use One Night Ultimate Werewolf (ONUW) as a testbed and present MultiMind, the first framework integrating multimodal information into SDG agents. MultiMind processes facial expressions and vocal tones alongside verbal content, while employing a Theory of Mind (ToM) model to represent each player's suspicion levels toward others. By combining this ToM model with Monte Carlo Tree Search (MCTS), our agent identifies communication strategies that minimize suspicion directed at itself. Through comprehensive evaluation in both agent-versus-agent simulations and studies with human players, we demonstrate MultiMind's superior performance in gameplay. Our work presents a significant advancement toward LLM agents capable of human-like social reasoning across multimodal domains.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18039
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
Zhang, Zheng
Xiao, Nuoqian
Chai, Qi
Ye, Deheng
Wang, Hao
Artificial Intelligence
Large Language Model (LLM) agents have demonstrated impressive capabilities in social deduction games (SDGs) like Werewolf, where strategic reasoning and social deception are essential. However, current approaches remain limited to textual information, ignoring crucial multimodal cues such as facial expressions and tone of voice that humans naturally use to communicate. Moreover, existing SDG agents primarily focus on inferring other players' identities without modeling how others perceive themselves or fellow players. To address these limitations, we use One Night Ultimate Werewolf (ONUW) as a testbed and present MultiMind, the first framework integrating multimodal information into SDG agents. MultiMind processes facial expressions and vocal tones alongside verbal content, while employing a Theory of Mind (ToM) model to represent each player's suspicion levels toward others. By combining this ToM model with Monte Carlo Tree Search (MCTS), our agent identifies communication strategies that minimize suspicion directed at itself. Through comprehensive evaluation in both agent-versus-agent simulations and studies with human players, we demonstrate MultiMind's superior performance in gameplay. Our work presents a significant advancement toward LLM agents capable of human-like social reasoning across multimodal domains.
title MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind
topic Artificial Intelligence
url https://arxiv.org/abs/2504.18039