Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Galisai, Marcello, Cifani, Susanna, Giarrusso, Francesco, Bisconti, Piercosma, Prandi, Matteo, Pierucci, Federico, Sartore, Federico, Nardi, Daniele |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025)
by: Prandi, Matteo, et al.
Published: (2025)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Metaphor Is Not All Attention Needs
by: Sorokoletova, Olga, et al.
Published: (2026)
by: Sorokoletova, Olga, et al.
Published: (2026)
Beyond Single-Agent Safety: A Taxonomy of Risks in LLM-to-LLM Interactions
by: Bisconti, Piercosma, et al.
Published: (2025)
by: Bisconti, Piercosma, et al.
Published: (2025)
Agentic Microphysics: A Manifesto for Generative AI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Institutional AI: A Governance Framework for Distributional AGI Safety
by: Pierucci, Federico, et al.
Published: (2026)
by: Pierucci, Federico, et al.
Published: (2026)
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot Markets via Public Governance Graphs
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
by: Syrnikov, Marcantonio Bracale, et al.
Published: (2026)
Standards for trustworthy AI in the European Union: technical rationale, structural challenges, and an implementation path
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Guarding the Guardrails: A Taxonomy-Driven Approach to Jailbreak Detection
by: Giarrusso, Francesco, et al.
Published: (2025)
by: Giarrusso, Francesco, et al.
Published: (2025)
A Participatory Strategy for AI Ethics in Education and Rehabilitation grounded in the Capability Approach
by: Cesaroni, Valeria, et al.
Published: (2025)
by: Cesaroni, Valeria, et al.
Published: (2025)
Adaptive Multimodal Agents-Based Framework for Automatic Workflow Execution
by: Cifani, Susanna, et al.
Published: (2026)
by: Cifani, Susanna, et al.
Published: (2026)
Learning from Mistakes: Can LLM Self-Recover after Misalignment?
by: Sorokoletova, Olga E., et al.
Published: (2026)
by: Sorokoletova, Olga E., et al.
Published: (2026)
Computer-assisted simultaneous interpreting
by: Prandi, Bianca
Published: (2023)
by: Prandi, Bianca
Published: (2023)
Interpretable Stylistic Variation in Human and LLM Writing Across Genres, Models, and Decoding Strategies
by: Rallapalli, Swati, et al.
Published: (2026)
by: Rallapalli, Swati, et al.
Published: (2026)
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
by: Riaz, Nauman, et al.
Published: (2024)
by: Riaz, Nauman, et al.
Published: (2024)
Detecting Stylistic Fingerprints of Large Language Models
by: Bitton, Yehonatan, et al.
Published: (2025)
by: Bitton, Yehonatan, et al.
Published: (2025)
A Multi-Perspective Benchmark and Moderation Model for Evaluating Safety and Adversarial Robustness
by: Machlovi, Naseem, et al.
Published: (2025)
by: Machlovi, Naseem, et al.
Published: (2025)
Zero-Shot Hierarchical Classification on the Common Procurement Vocabulary Taxonomy
by: Moiraghi, Federico, et al.
Published: (2024)
by: Moiraghi, Federico, et al.
Published: (2024)
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
Probing the Limits of Stylistic Alignment in Vision-Language Models
by: Farajidizaji, Asma, et al.
Published: (2025)
by: Farajidizaji, Asma, et al.
Published: (2025)
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
by: Chen, Sijia, et al.
Published: (2025)
by: Chen, Sijia, et al.
Published: (2025)
EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
by: Ma, Jicheng, et al.
Published: (2026)
by: Ma, Jicheng, et al.
Published: (2026)
Stylistic Evolution and LLM Neutrality in Singlish Language
by: Foo, Linus Tze En, et al.
Published: (2026)
by: Foo, Linus Tze En, et al.
Published: (2026)
Neurobiber: Fast and Interpretable Stylistic Feature Extraction
by: Alkiek, Kenan, et al.
Published: (2025)
by: Alkiek, Kenan, et al.
Published: (2025)
Are You Human? An Adversarial Benchmark to Expose LLMs
by: Gressel, Gilad, et al.
Published: (2024)
by: Gressel, Gilad, et al.
Published: (2024)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
StyleAdaptedLM: Enhancing Instruction Following Models with Efficient Stylistic Transfer
by: Ramu, Pritika, et al.
Published: (2025)
by: Ramu, Pritika, et al.
Published: (2025)
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
by: Bianchi, Federico, et al.
Published: (2023)
by: Bianchi, Federico, et al.
Published: (2023)
IDIOLEX: Unified and Continuous Representations for Idiolectal and Stylistic Variation
by: Kantharuban, Anjali, et al.
Published: (2026)
by: Kantharuban, Anjali, et al.
Published: (2026)
Internal Safety Collapse in Frontier Large Language Models
by: Wu, Yutao, et al.
Published: (2026)
by: Wu, Yutao, et al.
Published: (2026)
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
by: Zhu, Zihao, et al.
Published: (2025)
by: Zhu, Zihao, et al.
Published: (2025)
Efficient Uncertainty Estimation for LLM-based Entity Linking in Tabular Data
by: Bono, Carlo, et al.
Published: (2025)
by: Bono, Carlo, et al.
Published: (2025)
CLASE: A Hybrid Method for Chinese Legalese Stylistic Evaluation
by: Ma, Yiran Rex, et al.
Published: (2026)
by: Ma, Yiran Rex, et al.
Published: (2026)
Latent Space Interpretation for Stylistic Analysis and Explainable Authorship Attribution
by: Alshomary, Milad, et al.
Published: (2024)
by: Alshomary, Milad, et al.
Published: (2024)
Style over Substance: Distilled Language Models Reason Via Stylistic Replication
by: Lippmann, Philip, et al.
Published: (2025)
by: Lippmann, Philip, et al.
Published: (2025)
Stylistic Deceptions in Online News
by: Riggs, Ashley
Published: (2022)
by: Riggs, Ashley
Published: (2022)
Controllable Stylistic Text Generation with Train-Time Attribute-Regularized Diffusion
by: Zhou, Fan, et al.
Published: (2025)
by: Zhou, Fan, et al.
Published: (2025)
Similar Items
-
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
by: Bisconti, Piercosma, et al.
Published: (2025) -
Bench-2-CoP: Can We Trust Benchmarking for EU AI Compliance?
by: Prandi, Matteo, et al.
Published: (2025) -
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
by: Bisconti, Piercosma, et al.
Published: (2026) -
From Adversarial Poetry to Adversarial Tales: An Interpretability Research Agenda
by: Bisconti, Piercosma, et al.
Published: (2025) -
Metaphor Is Not All Attention Needs
by: Sorokoletova, Olga, et al.
Published: (2026)