On the Trade-Off Between Transparency and Security in Adversarial Machine Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fenaux, Lucas, Srinivasa, Christopher, Kerschbaum, Florian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917081047367680
author Fenaux, Lucas
Srinivasa, Christopher
Kerschbaum, Florian
author_facet Fenaux, Lucas
Srinivasa, Christopher
Kerschbaum, Florian
contents Transparency and security are both central to Responsible AI, but they may conflict in adversarial settings. We investigate the strategic effect of transparency for agents through the lens of transferable adversarial example attacks. In transferable adversarial example attacks, attackers maliciously perturb their inputs using surrogate models to fool a defender's target model. These models can be defended or undefended, with both players having to decide which to use. Using a large-scale empirical evaluation of nine attacks across 181 models, we find that attackers are more successful when they match the defender's decision; hence, obscurity could be beneficial to the defender. With game theory, we analyze this trade-off between transparency and security by modeling this problem as both a Nash game and a Stackelberg game, and comparing the expected outcomes. Our analysis confirms that only knowing whether a defender's model is defended or not can sometimes be enough to damage its security. This result serves as an indicator of the general trade-off between transparency and security, suggesting that transparency in AI systems can be at odds with security. Beyond adversarial machine learning, our work illustrates how game-theoretic reasoning can uncover conflicts between transparency and security.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11842
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On the Trade-Off Between Transparency and Security in Adversarial Machine Learning
Fenaux, Lucas
Srinivasa, Christopher
Kerschbaum, Florian
Machine Learning
Cryptography and Security
Computer Science and Game Theory
91A40 (Primary) 91A80, 68T05 (Secondary)
Transparency and security are both central to Responsible AI, but they may conflict in adversarial settings. We investigate the strategic effect of transparency for agents through the lens of transferable adversarial example attacks. In transferable adversarial example attacks, attackers maliciously perturb their inputs using surrogate models to fool a defender's target model. These models can be defended or undefended, with both players having to decide which to use. Using a large-scale empirical evaluation of nine attacks across 181 models, we find that attackers are more successful when they match the defender's decision; hence, obscurity could be beneficial to the defender. With game theory, we analyze this trade-off between transparency and security by modeling this problem as both a Nash game and a Stackelberg game, and comparing the expected outcomes. Our analysis confirms that only knowing whether a defender's model is defended or not can sometimes be enough to damage its security. This result serves as an indicator of the general trade-off between transparency and security, suggesting that transparency in AI systems can be at odds with security. Beyond adversarial machine learning, our work illustrates how game-theoretic reasoning can uncover conflicts between transparency and security.
title On the Trade-Off Between Transparency and Security in Adversarial Machine Learning
topic Machine Learning
Cryptography and Security
Computer Science and Game Theory
91A40 (Primary) 91A80, 68T05 (Secondary)
url https://arxiv.org/abs/2511.11842