Saved in:
Bibliographic Details
Main Authors: Yu, Xinquan, Lu, Wei, Luo, Xiangyang, Yang, Rui
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.02175
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911417536348160
author Yu, Xinquan
Lu, Wei
Luo, Xiangyang
Yang, Rui
author_facet Yu, Xinquan
Lu, Wei
Luo, Xiangyang
Yang, Rui
contents To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-grained annotations, such as patch/token-level annotations. This paper proposes a novel framework named Coupling Implicit and Explicit Cues (CIEC), which aims to achieve multimodal weakly-supervised manipulation localization for image-text pairs utilizing only coarse-grained image/sentence-level annotations. It comprises two branches, image-based and text-based weakly-supervised localization. For the former, we devise the Textual-guidance Refine Patch Selection (TRPS) module. It integrates forgery cues from both visual and textual perspectives to lock onto suspicious regions aided by spatial priors. Followed by the background silencing and spatial contrast constraints to suppress interference from irrelevant areas. For the latter, we devise the Visual-deviation Calibrated Token Grounding (VCTG) module. It focuses on meaningful content words and leverages relative visual bias to assist token localization. Followed by the asymmetric sparse and semantic consistency constraints to mitigate label noise and ensure reliability. Extensive experiments demonstrate the effectiveness of our CIEC, yielding results comparable to fully supervised methods on several evaluation metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02175
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization
Yu, Xinquan
Lu, Wei
Luo, Xiangyang
Yang, Rui
Computer Vision and Pattern Recognition
To mitigate the threat of misinformation, multimodal manipulation localization has garnered growing attention. Consider that current methods rely on costly and time-consuming fine-grained annotations, such as patch/token-level annotations. This paper proposes a novel framework named Coupling Implicit and Explicit Cues (CIEC), which aims to achieve multimodal weakly-supervised manipulation localization for image-text pairs utilizing only coarse-grained image/sentence-level annotations. It comprises two branches, image-based and text-based weakly-supervised localization. For the former, we devise the Textual-guidance Refine Patch Selection (TRPS) module. It integrates forgery cues from both visual and textual perspectives to lock onto suspicious regions aided by spatial priors. Followed by the background silencing and spatial contrast constraints to suppress interference from irrelevant areas. For the latter, we devise the Visual-deviation Calibrated Token Grounding (VCTG) module. It focuses on meaningful content words and leverages relative visual bias to assist token localization. Followed by the asymmetric sparse and semantic consistency constraints to mitigate label noise and ensure reliability. Extensive experiments demonstrate the effectiveness of our CIEC, yielding results comparable to fully supervised methods on several evaluation metrics.
title CIEC: Coupling Implicit and Explicit Cues for Multimodal Weakly Supervised Manipulation Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.02175