Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Panfilov, Alexander, Romov, Peter, Shilov, Igor, de Montjoye, Yves-Alexandre, Geiping, Jonas, Andriushchenko, Maksym
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911734057402368
author Panfilov, Alexander
Romov, Peter
Shilov, Igor
de Montjoye, Yves-Alexandre
Geiping, Jonas
Andriushchenko, Maksym
author_facet Panfilov, Alexander
Romov, Peter
Shilov, Igor
de Montjoye, Yves-Alexandre
Geiping, Jonas
Andriushchenko, Maksym
contents We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations. We deploy frontier agents, such as Claude Code and Codex, in an autoresearch loop with access to a library of 30+ prior methods and an evaluation script with a fixed compute budget. We show this pipeline to be effective in jailbreaking OpenAI's GPT-OSS-Safeguard-20B and in prompt injections against Meta-SecAlign-70B, an adversarially robust model. For GPT-OSS-Safeguard, the best agent-discovered method achieves up to 80\% attack success rate on CBRN queries, compared to <50\% for existing methods. For SecAlign, it achieves 100\% ASR, while the best prior automated methods only achieve 82\%. Notably, in our setting, attack methods are developed on unrelated surrogate models for a pure random-target token-forcing task, yet generalize directly to prompt injection on the adversarially trained model. Finally, we trace the lineage of methods developed during autoresearch, characterizing the agents' strategies and failure modes. Adversarial ML has long held that defenses must be evaluated against attacks tailored to them; autoresearch automates this principle, and we argue it should be the minimum bar for defense evaluation going forward.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24511
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Panfilov, Alexander
Romov, Peter
Shilov, Igor
de Montjoye, Yves-Alexandre
Geiping, Jonas
Andriushchenko, Maksym
Machine Learning
Artificial Intelligence
Cryptography and Security
We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injection evaluations. We deploy frontier agents, such as Claude Code and Codex, in an autoresearch loop with access to a library of 30+ prior methods and an evaluation script with a fixed compute budget. We show this pipeline to be effective in jailbreaking OpenAI's GPT-OSS-Safeguard-20B and in prompt injections against Meta-SecAlign-70B, an adversarially robust model. For GPT-OSS-Safeguard, the best agent-discovered method achieves up to 80\% attack success rate on CBRN queries, compared to <50\% for existing methods. For SecAlign, it achieves 100\% ASR, while the best prior automated methods only achieve 82\%. Notably, in our setting, attack methods are developed on unrelated surrogate models for a pure random-target token-forcing task, yet generalize directly to prompt injection on the adversarially trained model. Finally, we trace the lineage of methods developed during autoresearch, characterizing the agents' strategies and failure modes. Adversarial ML has long held that defenses must be evaluated against attacks tailored to them; autoresearch automates this principle, and we argue it should be the minimum bar for defense evaluation going forward.
title Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2603.24511