Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Galisai, Marcello, Cifani, Susanna, Giarrusso, Francesco, Bisconti, Piercosma, Prandi, Matteo, Pierucci, Federico, Sartore, Federico, Nardi, Daniele
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911608601575424
author Galisai, Marcello
Cifani, Susanna
Giarrusso, Francesco
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Sartore, Federico
Nardi, Daniele
author_facet Galisai, Marcello
Cifani, Susanna
Giarrusso, Francesco
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Sartore, Federico
Nardi, Daniele
contents The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety.
format Preprint
id arxiv_https___arxiv_org_abs_2604_18487
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
Galisai, Marcello
Cifani, Susanna
Giarrusso, Francesco
Bisconti, Piercosma
Prandi, Matteo
Pierucci, Federico
Sartore, Federico
Nardi, Daniele
Computation and Language
Artificial Intelligence
The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety.
title Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.18487