Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866916829392273408 |
|---|---|
| author | Le, Huy Hoan Nguyen, Van Sy Thinh Dang, Thi Le Chi Nguyen, Vo Thanh Khang Nguyen, Truong Thanh Hung Cao, Hung |
| author_facet | Le, Huy Hoan Nguyen, Van Sy Thinh Dang, Thi Le Chi Nguyen, Vo Thanh Khang Nguyen, Truong Thanh Hung Cao, Hung |
| contents | This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_04410 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models Le, Huy Hoan Nguyen, Van Sy Thinh Dang, Thi Le Chi Nguyen, Vo Thanh Khang Nguyen, Truong Thanh Hung Cao, Hung Computer Vision and Pattern Recognition Artificial Intelligence Information Retrieval I.2.10 This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios. |
| title | Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Information Retrieval I.2.10 |
| url | https://arxiv.org/abs/2507.04410 |