QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yuan, Ye, Song, Rui, Li, Weien, Li, Zeyu, Liu, Haochen, Kong, Xiangyu, Han, Changjiang, Yang, Yonghan, Zhao, Zichen, Dong, Zixuan, Lyu, Fuyuan, He, Bowei, Wu, Haolun, Kang, Jikun, Liu, Xue
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917536495304704
author Yuan, Ye
Song, Rui
Li, Weien
Li, Zeyu
Liu, Haochen
Kong, Xiangyu
Han, Changjiang
Yang, Yonghan
Zhao, Zichen
Dong, Zixuan
Lyu, Fuyuan
He, Bowei
Wu, Haolun
Kang, Jikun
Liu, Xue
author_facet Yuan, Ye
Song, Rui
Li, Weien
Li, Zeyu
Liu, Haochen
Kong, Xiangyu
Han, Changjiang
Yang, Yonghan
Zhao, Zichen
Dong, Zixuan
Lyu, Fuyuan
He, Bowei
Wu, Haolun
Kang, Jikun
Liu, Xue
contents Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction, making it difficult to tell whether an agent's language is actually grounded in what it perceived and did, or to identify the failure modes underlying its behavior. To address this gap, we introduce QUACK, an open-source environment and evaluation framework for auditing the grounding of agent language in multimodal social reasoning. QUACK evaluates agents at three levels: game outcomes, behavioral trajectories, and utterance-level consistency. Its core Statement Verification Pipeline reconstructs each agent's ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging spatial hallucination, unsupported accusation, deception collapse, and language-action inconsistency. Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, we find that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and makes over half of its accusations without grounded evidence. We release the full engine, evaluation framework, toolkit, and logs at https://github.com/AAAAA-Academia-Attractions/QUACK.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27068
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
Yuan, Ye
Song, Rui
Li, Weien
Li, Zeyu
Liu, Haochen
Kong, Xiangyu
Han, Changjiang
Yang, Yonghan
Zhao, Zichen
Dong, Zixuan
Lyu, Fuyuan
He, Bowei
Wu, Haolun
Kang, Jikun
Liu, Xue
Computation and Language
Artificial Intelligence
Multiagent Systems
Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and largely remain to text-only interaction, making it difficult to tell whether an agent's language is actually grounded in what it perceived and did, or to identify the failure modes underlying its behavior. To address this gap, we introduce QUACK, an open-source environment and evaluation framework for auditing the grounding of agent language in multimodal social reasoning. QUACK evaluates agents at three levels: game outcomes, behavioral trajectories, and utterance-level consistency. Its core Statement Verification Pipeline reconstructs each agent's ground-truth trajectory from engine logs and checks every discussion claim against it, automatically flagging spatial hallucination, unsupported accusation, deception collapse, and language-action inconsistency. Evaluating three frontier VLMs in both homogeneous and cross-model adversarial settings, we find that even the strongest agent hallucinates 15.1% of its verifiable spatial claims and makes over half of its accusations without grounded evidence. We release the full engine, evaluation framework, toolkit, and logs at https://github.com/AAAAA-Academia-Attractions/QUACK.
title QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents
topic Computation and Language
Artificial Intelligence
Multiagent Systems
url https://arxiv.org/abs/2605.27068