Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xuekang, Zhou, Ji-Zhe, Feng, Kaiwen, Qu, Chenfan, Wang, Xiwen, Wang, Yunfei, Zhou, Liting, Liu, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915928478842880
author Zhu, Xuekang
Zhou, Ji-Zhe
Feng, Kaiwen
Qu, Chenfan
Wang, Xiwen
Wang, Yunfei
Zhou, Liting
Liu, Jian
author_facet Zhu, Xuekang
Zhou, Ji-Zhe
Feng, Kaiwen
Qu, Chenfan
Wang, Xiwen
Wang, Yunfei
Zhou, Liting
Liu, Jian
contents With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipulation process, comprising a series of editing operations to create a deceptive image. However, existing IML methods remain manipulation-process-agnostic, directly producing localization masks in a one-shot prediction paradigm without modeling the underlying editing steps. This one-shot paradigm compresses the high-dimensional compositional space into a single binary mask, inducing severe dimensional collapse, which forces the model to discard essential structural cues and ultimately leads to overfitting and degraded generalization. To address this, we are the first to reformulate image manipulation localization as a conditional sequence prediction task, proposing the RITA framework. RITA predicts manipulated regions layer-by-layer in an ordered manner, using each step's prediction as the condition for the next, thereby explicitly modeling temporal dependencies and hierarchical structures among editing operations. To enable training and evaluation, we synthesize multi-step manipulation data and construct a new benchmark HSIM. We further propose the HSS metric to assess sequential order and hierarchical alignment. Extensive experiments show that: 1) RITA achieves SOTA generalization and robustness on traditional benchmarks; 2) it remains computationally efficient despite explicitly modeling multi-step sequences; and 3) it establishes a viable foundation for hierarchical, process-aware manipulation localization. Code and dataset are available at https://github.com/scu-zjz/RITA.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20006
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
Zhu, Xuekang
Zhou, Ji-Zhe
Feng, Kaiwen
Qu, Chenfan
Wang, Xiwen
Wang, Yunfei
Zhou, Liting
Liu, Jian
Computer Vision and Pattern Recognition
With the large models easing the labor-intensive manipulation process, image manipulations in today's real scenarios often entail a complex manipulation process, comprising a series of editing operations to create a deceptive image. However, existing IML methods remain manipulation-process-agnostic, directly producing localization masks in a one-shot prediction paradigm without modeling the underlying editing steps. This one-shot paradigm compresses the high-dimensional compositional space into a single binary mask, inducing severe dimensional collapse, which forces the model to discard essential structural cues and ultimately leads to overfitting and degraded generalization. To address this, we are the first to reformulate image manipulation localization as a conditional sequence prediction task, proposing the RITA framework. RITA predicts manipulated regions layer-by-layer in an ordered manner, using each step's prediction as the condition for the next, thereby explicitly modeling temporal dependencies and hierarchical structures among editing operations. To enable training and evaluation, we synthesize multi-step manipulation data and construct a new benchmark HSIM. We further propose the HSS metric to assess sequential order and hierarchical alignment. Extensive experiments show that: 1) RITA achieves SOTA generalization and robustness on traditional benchmarks; 2) it remains computationally efficient despite explicitly modeling multi-step sequences; and 3) it establishes a viable foundation for hierarchical, process-aware manipulation localization. Code and dataset are available at https://github.com/scu-zjz/RITA.
title Revisiting Image Manipulation Localization under Realistic Manipulation Scenarios
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.20006