SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Dai, Zeyu, Liu, Shengcai, He, Rui, Wu, Jiahao, Lu, Ning, Fan, Wenqi, Li, Qing, Tang, Ke
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912331112382464
author Dai, Zeyu
Liu, Shengcai
He, Rui
Wu, Jiahao
Lu, Ning
Fan, Wenqi
Li, Qing
Tang, Ke
author_facet Dai, Zeyu
Liu, Shengcai
He, Rui
Wu, Jiahao
Lu, Ning
Fan, Wenqi
Li, Qing
Tang, Ke
contents Unrestricted adversarial examples (UAEs), allow the attacker to create non-constrained adversarial examples without given clean samples, posing a severe threat to the safety of deep learning models. Recent works utilize diffusion models to generate UAEs. However, these UAEs often lack naturalness and imperceptibility due to simply optimizing in intermediate latent noises. In light of this, we propose SemDiff, a novel unrestricted adversarial attack that explores the semantic latent space of diffusion models for meaningful attributes, and devises a multi-attributes optimization approach to ensure attack success while maintaining the naturalness and imperceptibility of generated UAEs. We perform extensive experiments on four tasks on three high-resolution datasets, including CelebA-HQ, AFHQ and ImageNet. The results demonstrate that SemDiff outperforms state-of-the-art methods in terms of attack success rate and imperceptibility. The generated UAEs are natural and exhibit semantically meaningful changes, in accord with the attributes' weights. In addition, SemDiff is found capable of evading different defenses, which further validates its effectiveness and threatening.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11923
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models
Dai, Zeyu
Liu, Shengcai
He, Rui
Wu, Jiahao
Lu, Ning
Fan, Wenqi
Li, Qing
Tang, Ke
Machine Learning
Computer Vision and Pattern Recognition
Unrestricted adversarial examples (UAEs), allow the attacker to create non-constrained adversarial examples without given clean samples, posing a severe threat to the safety of deep learning models. Recent works utilize diffusion models to generate UAEs. However, these UAEs often lack naturalness and imperceptibility due to simply optimizing in intermediate latent noises. In light of this, we propose SemDiff, a novel unrestricted adversarial attack that explores the semantic latent space of diffusion models for meaningful attributes, and devises a multi-attributes optimization approach to ensure attack success while maintaining the naturalness and imperceptibility of generated UAEs. We perform extensive experiments on four tasks on three high-resolution datasets, including CelebA-HQ, AFHQ and ImageNet. The results demonstrate that SemDiff outperforms state-of-the-art methods in terms of attack success rate and imperceptibility. The generated UAEs are natural and exhibit semantically meaningful changes, in accord with the attributes' weights. In addition, SemDiff is found capable of evading different defenses, which further validates its effectiveness and threatening.
title SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.11923