Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Franzmeyer, Tim, McAleer, Stephen, Henriques, João F., Foerster, Jakob N., Torr, Philip H. S., Bibi, Adel, de Witt, Christian Schroeder
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911867046199296
author Franzmeyer, Tim
McAleer, Stephen
Henriques, João F.
Foerster, Jakob N.
Torr, Philip H. S.
Bibi, Adel
de Witt, Christian Schroeder
author_facet Franzmeyer, Tim
McAleer, Stephen
Henriques, João F.
Foerster, Jakob N.
Torr, Philip H. S.
Bibi, Adel
de Witt, Christian Schroeder
contents Autonomous agents deployed in the real world need to be robust against adversarial attacks on sensory inputs. Robustifying agent policies requires anticipating the strongest attacks possible. We demonstrate that existing observation-space attacks on reinforcement learning agents have a common weakness: while effective, their lack of information-theoretic detectability constraints makes them detectable using automated means or human inspection. Detectability is undesirable to adversaries as it may trigger security escalations. We introduce ε-illusory, a novel form of adversarial attack on sequential decision-makers that is both effective and of ε-bounded statistical detectability. We propose a novel dual ascent algorithm to learn such attacks end-to-end. Compared to existing attacks, we empirically find ε-illusory to be significantly harder to detect with automated methods, and a small study with human participants (IRB approval under reference R84123/RE001) suggests they are similarly harder to detect for humans. Our findings suggest the need for better anomaly detectors, as well as effective hardware- and system-level defenses. The project website can be found at https://tinyurl.com/illusory-attacks.
format Preprint
id arxiv_https___arxiv_org_abs_2207_10170
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks
Franzmeyer, Tim
McAleer, Stephen
Henriques, João F.
Foerster, Jakob N.
Torr, Philip H. S.
Bibi, Adel
de Witt, Christian Schroeder
Artificial Intelligence
Autonomous agents deployed in the real world need to be robust against adversarial attacks on sensory inputs. Robustifying agent policies requires anticipating the strongest attacks possible. We demonstrate that existing observation-space attacks on reinforcement learning agents have a common weakness: while effective, their lack of information-theoretic detectability constraints makes them detectable using automated means or human inspection. Detectability is undesirable to adversaries as it may trigger security escalations. We introduce ε-illusory, a novel form of adversarial attack on sequential decision-makers that is both effective and of ε-bounded statistical detectability. We propose a novel dual ascent algorithm to learn such attacks end-to-end. Compared to existing attacks, we empirically find ε-illusory to be significantly harder to detect with automated methods, and a small study with human participants (IRB approval under reference R84123/RE001) suggests they are similarly harder to detect for humans. Our findings suggest the need for better anomaly detectors, as well as effective hardware- and system-level defenses. The project website can be found at https://tinyurl.com/illusory-attacks.
title Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks
topic Artificial Intelligence
url https://arxiv.org/abs/2207.10170