Saved in:
Bibliographic Details
Main Authors: Chen, Xingbai, Fu, Tingchao, Liu, Renyang, Zhou, Wei, Yi, Chao
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.16157
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911167396446208
author Chen, Xingbai
Fu, Tingchao
Liu, Renyang
Zhou, Wei
Yi, Chao
author_facet Chen, Xingbai
Fu, Tingchao
Liu, Renyang
Zhou, Wei
Yi, Chao
contents Referring Expression Segmentation (RES) enables precise object segmentation in images based on natural language descriptions, offering high flexibility and broad applicability in real-world vision tasks. Despite its impressive performance, the robustness of RES models against adversarial examples remains largely unexplored. While prior adversarial attack methods have explored adversarial robustness on conventional segmentation models, they perform poorly when directly applied to RES models, failing to expose vulnerabilities in its multimodal structure. In practical open-world scenarios, users typically issue multiple, diverse referring expressions to interact with the same image, highlighting the need for adversarial examples that generalize across varied textual inputs. Furthermore, from the perspective of privacy protection, ensuring that RES models do not segment sensitive content without explicit authorization is a crucial aspect of enhancing the robustness and security of multimodal vision-language systems. To address these challenges, we present PEAT, an Embedding-Guided Bidirectional Attack for RES models. Extensive experiments across multiple RES architectures and standard benchmarks show that PEAT consistently outperforms competitive baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16157
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models
Chen, Xingbai
Fu, Tingchao
Liu, Renyang
Zhou, Wei
Yi, Chao
Computer Vision and Pattern Recognition
Referring Expression Segmentation (RES) enables precise object segmentation in images based on natural language descriptions, offering high flexibility and broad applicability in real-world vision tasks. Despite its impressive performance, the robustness of RES models against adversarial examples remains largely unexplored. While prior adversarial attack methods have explored adversarial robustness on conventional segmentation models, they perform poorly when directly applied to RES models, failing to expose vulnerabilities in its multimodal structure. In practical open-world scenarios, users typically issue multiple, diverse referring expressions to interact with the same image, highlighting the need for adversarial examples that generalize across varied textual inputs. Furthermore, from the perspective of privacy protection, ensuring that RES models do not segment sensitive content without explicit authorization is a crucial aspect of enhancing the robustness and security of multimodal vision-language systems. To address these challenges, we present PEAT, an Embedding-Guided Bidirectional Attack for RES models. Extensive experiments across multiple RES architectures and standard benchmarks show that PEAT consistently outperforms competitive baselines.
title Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.16157