AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Beyer, Tim, Dornbusch, Jonas, Steimle, Jakob, Ladenburger, Moritz, Schwinn, Leo, Günnemann, Stephan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911251340197888
author Beyer, Tim
Dornbusch, Jonas
Steimle, Jakob
Ladenburger, Moritz
Schwinn, Leo
Günnemann, Stephan
author_facet Beyer, Tim
Dornbusch, Jonas
Steimle, Jakob
Ladenburger, Moritz
Schwinn, Leo
Günnemann, Stephan
contents The rapid expansion of research on Large Language Model (LLM) safety and robustness has produced a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evaluation methods. This fragmentation makes reproducibility and comparability across studies challenging, hindering meaningful progress. To address these issues, we introduce AdversariaLLM, a toolbox for conducting LLM jailbreak robustness research. Its design centers on reproducibility, correctness, and extensibility. The framework implements twelve adversarial attack algorithms, integrates seven benchmark datasets spanning harmfulness, over-refusal, and utility evaluation, and provides access to a wide range of open-weight LLMs via Hugging Face. The implementation includes advanced features for comparability and reproducibility such as compute-resource tracking, deterministic results, and distributional evaluation techniques. \name also integrates judging through the companion package JudgeZoo, which can also be used independently. Together, these components aim to establish a robust foundation for transparent, comparable, and reproducible research in LLM safety.
format Preprint
id arxiv_https___arxiv_org_abs_2511_04316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
Beyer, Tim
Dornbusch, Jonas
Steimle, Jakob
Ladenburger, Moritz
Schwinn, Leo
Günnemann, Stephan
Artificial Intelligence
Software Engineering
The rapid expansion of research on Large Language Model (LLM) safety and robustness has produced a fragmented and oftentimes buggy ecosystem of implementations, datasets, and evaluation methods. This fragmentation makes reproducibility and comparability across studies challenging, hindering meaningful progress. To address these issues, we introduce AdversariaLLM, a toolbox for conducting LLM jailbreak robustness research. Its design centers on reproducibility, correctness, and extensibility. The framework implements twelve adversarial attack algorithms, integrates seven benchmark datasets spanning harmfulness, over-refusal, and utility evaluation, and provides access to a wide range of open-weight LLMs via Hugging Face. The implementation includes advanced features for comparability and reproducibility such as compute-resource tracking, deterministic results, and distributional evaluation techniques. \name also integrates judging through the companion package JudgeZoo, which can also be used independently. Together, these components aim to establish a robust foundation for transparent, comparable, and reproducible research in LLM safety.
title AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
topic Artificial Intelligence
Software Engineering
url https://arxiv.org/abs/2511.04316