RvB: Automating AI System Hardening via Iterative Red-Blue Games

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Lige, Liu, Zicheng, Zhang, Jie, Yan, Lewen, Liu, Dongrui, Shao, Jing
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914283965644800
author Huang, Lige
Liu, Zicheng
Zhang, Jie
Yan, Lewen
Liu, Dongrui
Shao, Jing
author_facet Huang, Lige
Liu, Zicheng
Zhang, Jie
Yan, Lewen
Liu, Dongrui
Shao, Jing
contents The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial adaptation hardening. To bridge this gap, we propose the Red Team vs. Blue Team (RvB) framework, formulated as a training-free, sequential, imperfect-information game. In this process, the Red Team exposes vulnerabilities, driving the Blue Team to learning effective solutions without parameter updates. We validate our framework across two challenging domains: dynamic code hardening against CVEs and guardrail optimization against jailbreaks. Our empirical results show that this interaction compels the Blue Team to learn fundamental defensive principles, leading to robust remediations that are not merely overfitted to specific exploits. RvB achieves Defense Success Rates of 90\% and 45\% across the respective tasks while maintaining near 0\% False Positive Rates, significantly surpassing baselines. This work establishes the iterative adversarial interaction framework as a practical paradigm that automates the continuous hardening of AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2601_19726
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RvB: Automating AI System Hardening via Iterative Red-Blue Games
Huang, Lige
Liu, Zicheng
Zhang, Jie
Yan, Lewen
Liu, Dongrui
Shao, Jing
Cryptography and Security
Artificial Intelligence
Computation and Language
The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial adaptation hardening. To bridge this gap, we propose the Red Team vs. Blue Team (RvB) framework, formulated as a training-free, sequential, imperfect-information game. In this process, the Red Team exposes vulnerabilities, driving the Blue Team to learning effective solutions without parameter updates. We validate our framework across two challenging domains: dynamic code hardening against CVEs and guardrail optimization against jailbreaks. Our empirical results show that this interaction compels the Blue Team to learn fundamental defensive principles, leading to robust remediations that are not merely overfitted to specific exploits. RvB achieves Defense Success Rates of 90\% and 45\% across the respective tasks while maintaining near 0\% False Positive Rates, significantly surpassing baselines. This work establishes the iterative adversarial interaction framework as a practical paradigm that automates the continuous hardening of AI systems.
title RvB: Automating AI System Hardening via Iterative Red-Blue Games
topic Cryptography and Security
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2601.19726