Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lian, Jingchun, Liu, Lingyu, Wang, Yaxiong, Wu, Yujiao, Wu, Lianwei, Zhu, Li, Zheng, Zhedong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917421164527616
author Lian, Jingchun
Liu, Lingyu
Wang, Yaxiong
Wu, Yujiao
Wu, Lianwei
Zhu, Li
Zheng, Zhedong
author_facet Lian, Jingchun
Liu, Lingyu
Wang, Yaxiong
Wu, Yujiao
Wu, Lianwei
Zhu, Li
Zheng, Zhedong
contents Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. To address this, we introduce Forgery Attribution Report Generation, a new multimodal task that jointly localizes forged regions ("Where") and generates natural language explanations grounded in the editing process ("Why"). This dual-focus approach goes beyond traditional forensics, providing a comprehensive understanding of the manipulation. To enable research in this domain, we present Multi-Modal Tamper Tracing (MMTT), a large-scale dataset of 152,217 samples, each with a process-derived ground-truth mask and a human-authored textual description, ensuring high annotation precision and linguistic richness. We further propose ForgeryTalker, a unified end-to-end framework that integrates vision and language via a shared encoder (image encoder + Q-former) and dual decoders for mask and text generation, enabling coherent cross-modal reasoning. Experiments show that ForgeryTalker achieves competitive performance on both report generation and forgery localization subtasks, i.e., 59.3 CIDEr and 73.67 IoU, respectively, establishing a baseline for explainable multimedia forensics. Dataset and code will be released to foster future research.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19685
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
Lian, Jingchun
Liu, Lingyu
Wang, Yaxiong
Wu, Yujiao
Wu, Lianwei
Zhu, Li
Zheng, Zhedong
Computer Vision and Pattern Recognition
Artificial Intelligence
Existing facial forgery detection methods typically focus on binary classification or pixel-level localization, providing little semantic insight into the nature of the manipulation. To address this, we introduce Forgery Attribution Report Generation, a new multimodal task that jointly localizes forged regions ("Where") and generates natural language explanations grounded in the editing process ("Why"). This dual-focus approach goes beyond traditional forensics, providing a comprehensive understanding of the manipulation. To enable research in this domain, we present Multi-Modal Tamper Tracing (MMTT), a large-scale dataset of 152,217 samples, each with a process-derived ground-truth mask and a human-authored textual description, ensuring high annotation precision and linguistic richness. We further propose ForgeryTalker, a unified end-to-end framework that integrates vision and language via a shared encoder (image encoder + Q-former) and dual decoders for mask and text generation, enabling coherent cross-modal reasoning. Experiments show that ForgeryTalker achieves competitive performance on both report generation and forgery localization subtasks, i.e., 59.3 CIDEr and 73.67 IoU, respectively, establishing a baseline for explainable multimedia forensics. Dataset and code will be released to foster future research.
title Generating Attribution Reports for Manipulated Facial Images: A Dataset and Baseline
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.19685