VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908842237886464 |
|---|---|
| author | Jung, Jaeyoon Yoon, Yejun Park, Kunwoo |
| author_facet | Jung, Jaeyoon Yoon, Yejun Park, Kunwoo |
| contents | This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple stages of fact-checking. Textual and visual evidence is retrieved from the knowledge store enriched through additional web collection. To identify key information and address inconsistencies among evidence items, modality-specific and cross-modal agents generate analysis reports. In the subsequent stage, question-answer pairs are produced based on these reports. Finally, the Verdict Prediction agent produces the verification outcome based on the image-text claim and the generated question-answer pairs. Our system ranked first on the leaderboard across all evaluation metrics. The source code is publicly available at https://github.com/ssu-humane/VILLAIN. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_04587 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration Jung, Jaeyoon Yoon, Yejun Park, Kunwoo Computation and Language Artificial Intelligence Computers and Society This paper describes VILLAIN, a multimodal fact-checking system that verifies image-text claims through prompt-based multi-agent collaboration. For the AVerImaTeC shared task, VILLAIN employs vision-language model agents across multiple stages of fact-checking. Textual and visual evidence is retrieved from the knowledge store enriched through additional web collection. To identify key information and address inconsistencies among evidence items, modality-specific and cross-modal agents generate analysis reports. In the subsequent stage, question-answer pairs are produced based on these reports. Finally, the Verdict Prediction agent produces the verification outcome based on the image-text claim and the generated question-answer pairs. Our system ranked first on the leaderboard across all evaluation metrics. The source code is publicly available at https://github.com/ssu-humane/VILLAIN. |
| title | VILLAIN at AVerImaTeC: Verifying Image-Text Claims via Multi-Agent Collaboration |
| topic | Computation and Language Artificial Intelligence Computers and Society |
| url | https://arxiv.org/abs/2602.04587 |