Model Mimic Attack: Knowledge Distillation for Provably Transferable Adversarial Examples

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lukyanov, Kirill, Perminov, Andrew, Turdakov, Denis, Pautov, Mikhail
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916447857410048
author Lukyanov, Kirill
Perminov, Andrew
Turdakov, Denis
Pautov, Mikhail
author_facet Lukyanov, Kirill
Perminov, Andrew
Turdakov, Denis
Pautov, Mikhail
contents The vulnerability of artificial neural networks to adversarial perturbations in the black-box setting is widely studied in the literature. The majority of attack methods to construct these perturbations suffer from an impractically large number of queries required to find an adversarial example. In this work, we focus on knowledge distillation as an approach to conduct transfer-based black-box adversarial attacks and propose an iterative training of the surrogate model on an expanding dataset. This work is the first, to our knowledge, to provide provable guarantees on the success of knowledge distillation-based attack on classification neural networks: we prove that if the student model has enough learning capabilities, the attack on the teacher model is guaranteed to be found within the finite number of distillation iterations.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15889
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Model Mimic Attack: Knowledge Distillation for Provably Transferable Adversarial Examples
Lukyanov, Kirill
Perminov, Andrew
Turdakov, Denis
Pautov, Mikhail
Machine Learning
Artificial Intelligence
The vulnerability of artificial neural networks to adversarial perturbations in the black-box setting is widely studied in the literature. The majority of attack methods to construct these perturbations suffer from an impractically large number of queries required to find an adversarial example. In this work, we focus on knowledge distillation as an approach to conduct transfer-based black-box adversarial attacks and propose an iterative training of the surrogate model on an expanding dataset. This work is the first, to our knowledge, to provide provable guarantees on the success of knowledge distillation-based attack on classification neural networks: we prove that if the student model has enough learning capabilities, the attack on the teacher model is guaranteed to be found within the finite number of distillation iterations.
title Model Mimic Attack: Knowledge Distillation for Provably Transferable Adversarial Examples
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2410.15889