RvB: Automating AI System Hardening via Iterative Red-Blue Games
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Lige, Liu, Zicheng, Zhang, Jie, Yan, Lewen, Liu, Dongrui, Shao, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025)
by: Liu, Zicheng, et al.
Published: (2025)
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025)
by: Hu, Xuhao, et al.
Published: (2025)
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024)
by: Jiang, Bojian, et al.
Published: (2024)
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
FSLH: Flexible Mechanized Speculative Load Hardening
by: Baumann, Jonathan, et al.
Published: (2025)
by: Baumann, Jonathan, et al.
Published: (2025)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models
by: Wei, Zhang, et al.
Published: (2025)
by: Wei, Zhang, et al.
Published: (2025)
Prompt Optimization and Evaluation for LLM Automated Red Teaming
by: Freenor, Michael, et al.
Published: (2025)
by: Freenor, Michael, et al.
Published: (2025)
Training a General Purpose Automated Red Teaming Model
by: Padmakumar, Aishwarya, et al.
Published: (2026)
by: Padmakumar, Aishwarya, et al.
Published: (2026)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction
by: Zhang, Jinchuan, et al.
Published: (2024)
by: Zhang, Jinchuan, et al.
Published: (2024)
RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning
by: Horal, Artur, et al.
Published: (2025)
by: Horal, Artur, et al.
Published: (2025)
Security Assessment and Hardening of Fog Computing Systems
by: Cesarano, Carmine
Published: (2023)
by: Cesarano, Carmine
Published: (2023)
PIG: Privacy Jailbreak Attack on LLMs via Gradient-based Iterative In-Context Optimization
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
by: Béjar, Mario Rodríguez, et al.
Published: (2026)
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability
by: Lu, Xiaoya, et al.
Published: (2025)
by: Lu, Xiaoya, et al.
Published: (2025)
Place Protections at the Right Place: Targeted Hardening for Cryptographic Code against Spectre v1
by: Zhu, Yiming, et al.
Published: (2024)
by: Zhu, Yiming, et al.
Published: (2024)
Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
by: Sahu, Anubhab, et al.
Published: (2026)
by: Sahu, Anubhab, et al.
Published: (2026)
Resource Consumption Red-Teaming for Large Vision-Language Models
by: Gao, Haoran, et al.
Published: (2025)
by: Gao, Haoran, et al.
Published: (2025)
Automated Formal Verification of a Software Fault Isolation System
by: Sotoudeh, Matthew, et al.
Published: (2025)
by: Sotoudeh, Matthew, et al.
Published: (2025)
CoTGuard: Using Chain-of-Thought Triggering for Copyright Protection in Multi-Agent LLM Systems
by: Wen, Yan, et al.
Published: (2025)
by: Wen, Yan, et al.
Published: (2025)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
by: Yuan, Xiaohan, et al.
Published: (2024)
by: Yuan, Xiaohan, et al.
Published: (2024)
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
by: Xie, Yuanbo, et al.
Published: (2026)
by: Xie, Yuanbo, et al.
Published: (2026)
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
by: Yarmohammadtoosky, Sahar, et al.
Published: (2025)
PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
by: Munoz, Gary D. Lopez, et al.
Published: (2024)
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
by: Zhang, Xiaozhe, et al.
Published: (2026)
by: Zhang, Xiaozhe, et al.
Published: (2026)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
by: Li, Lijun, et al.
Published: (2025)
by: Li, Lijun, et al.
Published: (2025)
MIRAGE: Multimodal Immersive Reasoning and Guided Exploration for Red-Team Jailbreak Attacks
by: You, Wenhao, et al.
Published: (2025)
by: You, Wenhao, et al.
Published: (2025)
LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems
by: Zhang, Yuanhe, et al.
Published: (2025)
by: Zhang, Yuanhe, et al.
Published: (2025)
MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
iResolveX: Multi-Layered Indirect Call Resolution via Static Reasoning and Learning-Augmented Refinement
by: Santra, Monika, et al.
Published: (2026)
by: Santra, Monika, et al.
Published: (2026)
AJAR: Adaptive Jailbreak Architecture for Red-teaming
by: Dou, Yipu, et al.
Published: (2026)
by: Dou, Yipu, et al.
Published: (2026)
Large Language Models for Code: Security Hardening and Adversarial Testing
by: He, Jingxuan, et al.
Published: (2023)
by: He, Jingxuan, et al.
Published: (2023)
A Red Teaming Framework for Evaluating Robustness of AI-enabled Security Orchestration, Automation, and Response Systems
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
by: Shaikh, Ayan Javeed, et al.
Published: (2026)
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
by: Yang, Wenkui, et al.
Published: (2026)
by: Yang, Wenkui, et al.
Published: (2026)
Proactive Hardening of LLM Defenses with HASTE
by: Chen, Henry, et al.
Published: (2026)
by: Chen, Henry, et al.
Published: (2026)
AdvAgent: Controllable Blackbox Red-teaming on Web Agents
by: Xu, Chejian, et al.
Published: (2024)
by: Xu, Chejian, et al.
Published: (2024)
From Paranoia to Compliance: The Bumpy Road of System Hardening Practices on Stack Exchange
by: Busch, Niklas, et al.
Published: (2025)
by: Busch, Niklas, et al.
Published: (2025)
Similar Items
-
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
by: Liu, Zicheng, et al.
Published: (2025) -
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
by: Hu, Xuhao, et al.
Published: (2025) -
Automated Progressive Red Teaming
by: Jiang, Bojian, et al.
Published: (2024) -
REEF: Representation Encoding Fingerprints for Large Language Models
by: Zhang, Jie, et al.
Published: (2024) -
FSLH: Flexible Mechanized Speculative Load Hardening
by: Baumann, Jonathan, et al.
Published: (2025)