Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Le, Huy Hoan, Nguyen, Van Sy Thinh, Dang, Thi Le Chi, Nguyen, Vo Thanh Khang, Nguyen, Truong Thanh Hung, Cao, Hung
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916829392273408
author Le, Huy Hoan
Nguyen, Van Sy Thinh
Dang, Thi Le Chi
Nguyen, Vo Thanh Khang
Nguyen, Truong Thanh Hung
Cao, Hung
author_facet Le, Huy Hoan
Nguyen, Van Sy Thinh
Dang, Thi Le Chi
Nguyen, Vo Thanh Khang
Nguyen, Truong Thanh Hung
Cao, Hung
contents This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2507_04410
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
Le, Huy Hoan
Nguyen, Van Sy Thinh
Dang, Thi Le Chi
Nguyen, Vo Thanh Khang
Nguyen, Truong Thanh Hung
Cao, Hung
Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
I.2.10
This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios.
title Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Information Retrieval
I.2.10
url https://arxiv.org/abs/2507.04410