VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jung, Jaeyoon, Yoon, Yejun, Park, Kunwoo
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908842237886464
author Jung, Jaeyoon
Yoon, Yejun
Park, Kunwoo
author_facet Jung, Jaeyoon
Yoon, Yejun
Park, Kunwoo
contents This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple stages of fact-checking. Textual and visual evidence is retrieved from the knowledge store enriched through additional web collection. To identify key information and address inconsistencies among evidence items, modality-specific and cross-modal agents generate analysis reports. In the subsequent stage, question-answer pairs are produced based on these reports. Finally, the Verdict Prediction agent produces the verification outcome based on the image-text claim and the generated question-answer pairs. Our system ranked first on the leaderboard across all evaluation metrics. The source code is publicly available at https://github.com/ssu-humane/VILLAIN.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04587
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
Jung, Jaeyoon
Yoon, Yejun
Park, Kunwoo
Computation and Language
Artificial Intelligence
Computers and Society
This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple stages of fact-checking. Textual and visual evidence is retrieved from the knowledge store enriched through additional web collection. To identify key information and address inconsistencies among evidence items, modality-specific and cross-modal agents generate analysis reports. In the subsequent stage, question-answer pairs are produced based on these reports. Finally, the Verdict Prediction agent produces the verification outcome based on the image-text claim and the generated question-answer pairs. Our system ranked first on the leaderboard across all evaluation metrics. The source code is publicly available at https://github.com/ssu-humane/VILLAIN.
title VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
topic Computation and Language
Artificial Intelligence
Computers and Society
url https://arxiv.org/abs/2602.04587