Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sarkar, Soumyendu, Babu, Ashwin Ramesh, Mousavi, Sajad, Gundecha, Vineet, Ghorbanpour, Sahand, Naug, Avisek, Gutierrez, Ricardo Luna, Guillen, Antonio
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916690297618432
author Sarkar, Soumyendu
Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Gutierrez, Ricardo Luna
Guillen, Antonio
author_facet Sarkar, Soumyendu
Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Gutierrez, Ricardo Luna
Guillen, Antonio
contents We present a Reinforcement Learning Platform for Adversarial Black-box untargeted and targeted attacks, RLAB, that allows users to select from various distortion filters to create adversarial examples. The platform uses a Reinforcement Learning agent to add minimum distortion to input images while still causing misclassification by the target model. The agent uses a novel dual-action method to explore the input image at each step to identify sensitive regions for adding distortions while removing noises that have less impact on the target model. This dual action leads to faster and more efficient convergence of the attack. The platform can also be used to measure the robustness of image classification models against specific distortion types. Also, retraining the model with adversarial samples significantly improved robustness when evaluated on benchmark datasets. The proposed platform outperforms state-of-the-art methods in terms of the average number of queries required to cause misclassification. This advances trustworthiness with a positive social impact.
format Preprint
id arxiv_https___arxiv_org_abs_2501_14122
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
Sarkar, Soumyendu
Babu, Ashwin Ramesh
Mousavi, Sajad
Gundecha, Vineet
Ghorbanpour, Sahand
Naug, Avisek
Gutierrez, Ricardo Luna
Guillen, Antonio
Machine Learning
Artificial Intelligence
Cryptography and Security
Computer Vision and Pattern Recognition
We present a Reinforcement Learning Platform for Adversarial Black-box untargeted and targeted attacks, RLAB, that allows users to select from various distortion filters to create adversarial examples. The platform uses a Reinforcement Learning agent to add minimum distortion to input images while still causing misclassification by the target model. The agent uses a novel dual-action method to explore the input image at each step to identify sensitive regions for adding distortions while removing noises that have less impact on the target model. This dual action leads to faster and more efficient convergence of the attack. The platform can also be used to measure the robustness of image classification models against specific distortion types. Also, retraining the model with adversarial samples significantly improved robustness when evaluated on benchmark datasets. The proposed platform outperforms state-of-the-art methods in terms of the average number of queries required to cause misclassification. This advances trustworthiness with a positive social impact.
title Reinforcement Learning Platform for Adversarial Black-box Attacks with Custom Distortion Filters
topic Machine Learning
Artificial Intelligence
Cryptography and Security
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.14122