Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Swaminathan, Arjhun, Akgün, Mete
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915948903006208
author Swaminathan, Arjhun
Akgün, Mete
author_facet Swaminathan, Arjhun
Akgün, Mete
contents Deep neural networks for image classification remain vulnerable to adversarial examples -- small, imperceptible perturbations that induce misclassifications. In black-box settings, where only the final prediction is accessible, crafting targeted attacks that aim to misclassify into a specific target class is particularly challenging due to narrow decision regions. Current state-of-the-art methods often exploit the geometric properties of the decision boundary separating a source image and a target image rather than incorporating information from the images themselves. In contrast, we propose Targeted Edge-informed Attack (TEA), a novel attack that utilizes edge information from the target image to carefully perturb it, thereby producing an adversarial image that is closer to the source image while still achieving the desired target classification. Our approach consistently outperforms current state-of-the-art methods across different models in low query settings (nearly 70% fewer queries are used), a scenario especially relevant in real-world applications with limited queries and black-box access. Furthermore, by efficiently generating a suitable adversarial example, TEA provides an improved target initialization for established geometry-based attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2505_16313
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
Swaminathan, Arjhun
Akgün, Mete
Computer Vision and Pattern Recognition
Machine Learning
Deep neural networks for image classification remain vulnerable to adversarial examples -- small, imperceptible perturbations that induce misclassifications. In black-box settings, where only the final prediction is accessible, crafting targeted attacks that aim to misclassify into a specific target class is particularly challenging due to narrow decision regions. Current state-of-the-art methods often exploit the geometric properties of the decision boundary separating a source image and a target image rather than incorporating information from the images themselves. In contrast, we propose Targeted Edge-informed Attack (TEA), a novel attack that utilizes edge information from the target image to carefully perturb it, thereby producing an adversarial image that is closer to the source image while still achieving the desired target classification. Our approach consistently outperforms current state-of-the-art methods across different models in low query settings (nearly 70% fewer queries are used), a scenario especially relevant in real-world applications with limited queries and black-box access. Furthermore, by efficiently generating a suitable adversarial example, TEA provides an improved target initialization for established geometry-based attacks.
title Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2505.16313