LogoStyleFool: Vitiating Video Recognition Systems via Logo Style Transfer

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Cao, Yuxin, Zhao, Ziyu, Xiao, Xi, Wang, Derui, Xue, Minhui, Lu, Jin
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914735330426880
author Cao, Yuxin
Zhao, Ziyu
Xiao, Xi
Wang, Derui
Xue, Minhui
Lu, Jin
author_facet Cao, Yuxin
Zhao, Ziyu
Xiao, Xi
Wang, Derui
Xue, Minhui
Lu, Jin
contents Video recognition systems are vulnerable to adversarial examples. Recent studies show that style transfer-based and patch-based unrestricted perturbations can effectively improve attack efficiency. These attacks, however, face two main challenges: 1) Adding large stylized perturbations to all pixels reduces the naturalness of the video and such perturbations can be easily detected. 2) Patch-based video attacks are not extensible to targeted attacks due to the limited search space of reinforcement learning that has been widely used in video attacks recently. In this paper, we focus on the video black-box setting and propose a novel attack framework named LogoStyleFool by adding a stylized logo to the clean video. We separate the attack into three stages: style reference selection, reinforcement-learning-based logo style transfer, and perturbation optimization. We solve the first challenge by scaling down the perturbation range to a regional logo, while the second challenge is addressed by complementing an optimization stage after reinforcement learning. Experimental results substantiate the overall superiority of LogoStyleFool over three state-of-the-art patch-based attacks in terms of attack performance and semantic preservation. Meanwhile, LogoStyleFool still maintains its performance against two existing patch-based defense methods. We believe that our research is beneficial in increasing the attention of the security community to such subregional style transfer attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09935
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle LogoStyleFool: Vitiating Video Recognition Systems via Logo Style Transfer
Cao, Yuxin
Zhao, Ziyu
Xiao, Xi
Wang, Derui
Xue, Minhui
Lu, Jin
Computer Vision and Pattern Recognition
Cryptography and Security
Video recognition systems are vulnerable to adversarial examples. Recent studies show that style transfer-based and patch-based unrestricted perturbations can effectively improve attack efficiency. These attacks, however, face two main challenges: 1) Adding large stylized perturbations to all pixels reduces the naturalness of the video and such perturbations can be easily detected. 2) Patch-based video attacks are not extensible to targeted attacks due to the limited search space of reinforcement learning that has been widely used in video attacks recently. In this paper, we focus on the video black-box setting and propose a novel attack framework named LogoStyleFool by adding a stylized logo to the clean video. We separate the attack into three stages: style reference selection, reinforcement-learning-based logo style transfer, and perturbation optimization. We solve the first challenge by scaling down the perturbation range to a regional logo, while the second challenge is addressed by complementing an optimization stage after reinforcement learning. Experimental results substantiate the overall superiority of LogoStyleFool over three state-of-the-art patch-based attacks in terms of attack performance and semantic preservation. Meanwhile, LogoStyleFool still maintains its performance against two existing patch-based defense methods. We believe that our research is beneficial in increasing the attention of the security community to such subregional style transfer attacks.
title LogoStyleFool: Vitiating Video Recognition Systems via Logo Style Transfer
topic Computer Vision and Pattern Recognition
Cryptography and Security
url https://arxiv.org/abs/2312.09935