MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lu, Weihai, Zhao, Zhejun, Li, Yanshu, He, Huan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914521173458944
author Lu, Weihai
Zhao, Zhejun
Li, Yanshu
He, Huan
author_facet Lu, Weihai
Zhao, Zhejun
Li, Yanshu
He, Huan
contents Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual grounding, cross-modal interpretation ambiguity, and single-pass reasoning fragility. To address these, we propose Retrieval-Augmented Multi-modal Multi-agent Stance Detection (MM-StanceDet), a novel multi-agent framework integrating Retrieval Augmentation for contextual grounding, specialized Multimodal Analysis agents for nuanced interpretation, a Reasoning-Enhanced Debate stage for exploring perspectives, and Self-Reflection for robust adjudication. Extensive experiments on five datasets demonstrate MM-StanceDet significantly outperforms state-of-the-art baselines, validating the efficacy of its multi-agent architecture and structured reasoning stages in addressing complex multimodal stance challenges.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27934
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
Lu, Weihai
Zhao, Zhejun
Li, Yanshu
He, Huan
Artificial Intelligence
Computation and Language
Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual grounding, cross-modal interpretation ambiguity, and single-pass reasoning fragility. To address these, we propose Retrieval-Augmented Multi-modal Multi-agent Stance Detection (MM-StanceDet), a novel multi-agent framework integrating Retrieval Augmentation for contextual grounding, specialized Multimodal Analysis agents for nuanced interpretation, a Reasoning-Enhanced Debate stage for exploring perspectives, and Self-Reflection for robust adjudication. Extensive experiments on five datasets demonstrate MM-StanceDet significantly outperforms state-of-the-art baselines, validating the efficacy of its multi-agent architecture and structured reasoning stages in addressing complex multimodal stance challenges.
title MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
topic Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.27934