Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zongmin, Sun, Zhen, Liao, Yifan, Dong, Wenhan, He, Xinlei, Han, Xingshuo, Xu, Shengmin, Huang, Xinyi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913074543329280
author Zhang, Zongmin
Sun, Zhen
Liao, Yifan
Dong, Wenhan
He, Xinlei
Han, Xingshuo
Xu, Shengmin
Huang, Xinyi
author_facet Zhang, Zongmin
Sun, Zhen
Liao, Yifan
Dong, Wenhan
He, Xinlei
Han, Xingshuo
Xu, Shengmin
Huang, Xinyi
contents Prompt-driven Video Segmentation Foundation Models (VSFMs), such as SAM2, are increasingly used in applications including autonomous driving and digital pathology, yet their security risks remain underexplored. We study backdoor attacks against VSFMs and show that directly applying classic attacks such as BadNet is largely ineffective, yielding attack success rates (ASR) below 5%. Through gradient-similarity and attention-map analyses, we find that traditional backdoor training fails because clean and triggered samples induce aligned image-encoder gradients, while model attention remains focused on the prompt-specified object rather than the trigger. To address this limitation, we propose BadVSFM, the first backdoor attack framework tailored to prompt-driven VSFMs. BadVSFM uses a two-stage strategy that first learns trigger-specific encoder features and then trains the decoder to map triggered frame prompt representations to an attacker-specified target mask while preserving clean segmentation behavior. Experiments on five VSFMs and two datasets show that BadVSFM achieves strong, controllable backdoor effects across triggers and prompt types with limited clean-performance degradation. Ablations and interpretability analyses validate the necessity of the two-stage design, and five representative defenses remain largely ineffective. Our results reveal a practical and underexplored vulnerability of current VSFMs to backdoor threats.
format Preprint
id arxiv_https___arxiv_org_abs_2512_22046
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
Zhang, Zongmin
Sun, Zhen
Liao, Yifan
Dong, Wenhan
He, Xinlei
Han, Xingshuo
Xu, Shengmin
Huang, Xinyi
Computer Vision and Pattern Recognition
Cryptography and Security
Prompt-driven Video Segmentation Foundation Models (VSFMs), such as SAM2, are increasingly used in applications including autonomous driving and digital pathology, yet their security risks remain underexplored. We study backdoor attacks against VSFMs and show that directly applying classic attacks such as BadNet is largely ineffective, yielding attack success rates (ASR) below 5%. Through gradient-similarity and attention-map analyses, we find that traditional backdoor training fails because clean and triggered samples induce aligned image-encoder gradients, while model attention remains focused on the prompt-specified object rather than the trigger. To address this limitation, we propose BadVSFM, the first backdoor attack framework tailored to prompt-driven VSFMs. BadVSFM uses a two-stage strategy that first learns trigger-specific encoder features and then trains the decoder to map triggered frame prompt representations to an attacker-specified target mask while preserving clean segmentation behavior. Experiments on five VSFMs and two datasets show that BadVSFM achieves strong, controllable backdoor effects across triggers and prompt types with limited clean-performance degradation. Ablations and interpretability analyses validate the necessity of the two-stage design, and five representative defenses remain largely ineffective. Our results reveal a practical and underexplored vulnerability of current VSFMs to backdoor threats.
title Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2512.22046