AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
Fuente:
arXiv
Saved in:
| Main Authors: | Beyer, Tim, Dornbusch, Jonas, Steimle, Jakob, Ladenburger, Moritz, Schwinn, Leo, Günnemann, Stephan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
Fast Proxies for LLM Robustness Evaluation
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Automated Machine Learning: A Case Study on Non-Intrusive Appliance Load Monitoring
by: Moin, Armin, et al.
Published: (2022)
by: Moin, Armin, et al.
Published: (2022)
Prompt2DAG: A Modular Methodology for LLM-Based Data Enrichment Pipeline Generation
by: Alidu, Abubakari, et al.
Published: (2025)
by: Alidu, Abubakari, et al.
Published: (2025)
LLM-Based Robustness Testing of Microservice Applications: An Empirical Study
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
by: Tigulla, Hrushitha Goud, et al.
Published: (2026)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
by: Imran, Mia Mohammad, et al.
Published: (2025)
by: Imran, Mia Mohammad, et al.
Published: (2025)
MLE-Toolbox: An Open-Source Toolbox for Comprehensive EEG and MEG Data Analysis
by: Liu, Xiaobo
Published: (2026)
by: Liu, Xiaobo
Published: (2026)
ScriptSmith: A Unified LLM Framework for Enhancing IT Operations via Automated Bash Script Generation, Assessment, and Refinement
by: Chatterjee, Oishik, et al.
Published: (2024)
by: Chatterjee, Oishik, et al.
Published: (2024)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
RESCORE: LLM-Driven Simulation Recovery in Control Systems Research Papers
by: Bhat, Vineet, et al.
Published: (2026)
by: Bhat, Vineet, et al.
Published: (2026)
From Defects to Demands: A Unified, Iterative, and Heuristically Guided LLM-Based Framework for Automated Software Repair and Requirement Realization
by: Alex, et al.
Published: (2024)
by: Alex, et al.
Published: (2024)
Learn to Code Sustainably: An Empirical Study on LLM-based Green Code Generation
by: Vartziotis, Tina, et al.
Published: (2024)
by: Vartziotis, Tina, et al.
Published: (2024)
LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research
by: Yan, Shuo, et al.
Published: (2025)
by: Yan, Shuo, et al.
Published: (2025)
AutoEmpirical: LLM-Based Automated Research for Empirical Software Fault Analysis
by: Yu, Jiongchi, et al.
Published: (2025)
by: Yu, Jiongchi, et al.
Published: (2025)
LLM-Rosetta: A Hub-and-Spoke Intermediate Representation for Cross-Provider LLM API Translation
by: Ding, Peng
Published: (2026)
by: Ding, Peng
Published: (2026)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
by: Zhao, Shanhui, et al.
Published: (2025)
by: Zhao, Shanhui, et al.
Published: (2025)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
by: Yang, Ruozhao, et al.
Published: (2025)
by: Yang, Ruozhao, et al.
Published: (2025)
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
by: Liu, Jiacheng, et al.
Published: (2024)
by: Liu, Jiacheng, et al.
Published: (2024)
Harden and Catch for Just-in-Time Assured LLM-Based Software Testing: Open Research Challenges
by: Harman, Mark, et al.
Published: (2025)
by: Harman, Mark, et al.
Published: (2025)
A Performance Study of LLM-Generated Code on Leetcode
by: Coignion, Tristan, et al.
Published: (2024)
by: Coignion, Tristan, et al.
Published: (2024)
Specification and Detection of LLM Code Smells
by: Mahmoudi, Brahim, et al.
Published: (2025)
by: Mahmoudi, Brahim, et al.
Published: (2025)
Evaluating the effectiveness of LLM-based interoperability
by: Falcão, Rodrigo, et al.
Published: (2025)
by: Falcão, Rodrigo, et al.
Published: (2025)
Investigating The Smells of LLM Generated Code
by: Paul, Debalina Ghosh, et al.
Published: (2025)
by: Paul, Debalina Ghosh, et al.
Published: (2025)
Uncertainty Propagation in LLM-Based Systems
by: Xia, Boming, et al.
Published: (2026)
by: Xia, Boming, et al.
Published: (2026)
Breaking the Illusion of Identity in LLM Tooling
by: Miller, Marek
Published: (2026)
by: Miller, Marek
Published: (2026)
Reducing Cost of LLM Agents with Trajectory Reduction
by: Xiao, Yuan-An, et al.
Published: (2025)
by: Xiao, Yuan-An, et al.
Published: (2025)
LLM Collaboration With Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2025)
by: Liu, Shuo, et al.
Published: (2025)
Uncertainty Quantification for LLM-based Code Generation
by: Xu, Senrong, et al.
Published: (2026)
by: Xu, Senrong, et al.
Published: (2026)
Performance Review on LLM for solving leetcode problems
by: Wang, Lun, et al.
Published: (2025)
by: Wang, Lun, et al.
Published: (2025)
LLM-based Iterative Approach to Metamodeling in Automotive
by: Petrovic, Nenad, et al.
Published: (2025)
by: Petrovic, Nenad, et al.
Published: (2025)
Impact of Comments on LLM Comprehension of Legacy Code
by: Sabetto, Rock, et al.
Published: (2025)
by: Sabetto, Rock, et al.
Published: (2025)
Rover: Context-aware Conflict Resolution with LLM
by: Zhang, Qingyu, et al.
Published: (2026)
by: Zhang, Qingyu, et al.
Published: (2026)
LLM Applications: Current Paradigms and the Next Frontier
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
Strategic Decision Framework for Enterprise LLM Adoption
by: Trusov, Michael, et al.
Published: (2025)
by: Trusov, Michael, et al.
Published: (2025)
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning
by: Sun, Yuqiang, et al.
Published: (2024)
by: Sun, Yuqiang, et al.
Published: (2024)
VeriDebug: A Unified LLM for Verilog Debugging via Contrastive Embedding and Guided Correction
by: Wang, Ning, et al.
Published: (2025)
by: Wang, Ning, et al.
Published: (2025)
LLM Code Customization with Visual Results: A Benchmark on TikZ
by: Reux, Charly, et al.
Published: (2025)
by: Reux, Charly, et al.
Published: (2025)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
by: Jeong, Cheonsu, et al.
Published: (2026)
by: Jeong, Cheonsu, et al.
Published: (2026)
Similar Items
-
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026) -
Fast Proxies for LLM Robustness Evaluation
by: Beyer, Tim, et al.
Published: (2025) -
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025) -
Automated Machine Learning: A Case Study on Non-Intrusive Appliance Load Monitoring
by: Moin, Armin, et al.
Published: (2022) -
Prompt2DAG: A Modular Methodology for LLM-Based Data Enrichment Pipeline Generation
by: Alidu, Abubakari, et al.
Published: (2025)