Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866911608601575424 |
|---|---|
| author | Galisai, Marcello Cifani, Susanna Giarrusso, Francesco Bisconti, Piercosma Prandi, Matteo Pierucci, Federico Sartore, Federico Nardi, Daniele |
| author_facet | Galisai, Marcello Cifani, Susanna Giarrusso, Francesco Bisconti, Piercosma Prandi, Matteo Pierucci, Federico Sartore, Federico Nardi, Daniele |
| contents | The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_18487 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety Galisai, Marcello Cifani, Susanna Giarrusso, Francesco Bisconti, Piercosma Prandi, Matteo Pierucci, Federico Sartore, Federico Nardi, Daniele Computation and Language Artificial Intelligence The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety. |
| title | Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2604.18487 |