Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
Fuente:
arXiv
Saved in:
| Main Authors: | Schwinn, Leo, Scholten, Yan, Wollschläger, Tom, Xhonneux, Sophie, Casper, Stephen, Günnemann, Stephan, Gidel, Gauthier |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024)
by: Schwinn, Leo, et al.
Published: (2024)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025)
by: Scholten, Yan, et al.
Published: (2025)
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
Diffusion LLMs are Natural Adversaries for any LLM
by: Lüdke, David, et al.
Published: (2025)
by: Lüdke, David, et al.
Published: (2025)
Sampling-aware Adversarial Attacks Against Large Language Models
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
by: Campi, Francesco, et al.
Published: (2023)
by: Campi, Francesco, et al.
Published: (2023)
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More
by: Schuchardt, Jan, et al.
Published: (2023)
by: Schuchardt, Jan, et al.
Published: (2023)
Assessing Robustness via Score-Based Adversarial Image Generation
by: Kollovieh, Marcel, et al.
Published: (2023)
by: Kollovieh, Marcel, et al.
Published: (2023)
Adversarial Robustness of Graph Transformers
by: Foth, Philipp, et al.
Published: (2024)
by: Foth, Philipp, et al.
Published: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
In-Context Learning Can Re-learn Forbidden Tasks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
Efficient Time Series Processing for Transformers and State-Space Models through Token Merging
by: Götz, Leon, et al.
Published: (2024)
by: Götz, Leon, et al.
Published: (2024)
Byte Pair Encoding for Efficient Time Series Forecasting
by: Götz, Leon, et al.
Published: (2025)
by: Götz, Leon, et al.
Published: (2025)
What Expressivity Theory Misses: Message Passing Complexity for GNNs
by: Kemper, Niklas, et al.
Published: (2025)
by: Kemper, Niklas, et al.
Published: (2025)
Energy-based Epistemic Uncertainty for Graph Neural Networks
by: Fuchsgruber, Dominik, et al.
Published: (2024)
by: Fuchsgruber, Dominik, et al.
Published: (2024)
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
by: Geisler, Simon, et al.
Published: (2025)
by: Geisler, Simon, et al.
Published: (2025)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
by: Tomov, Tim, et al.
Published: (2025)
by: Tomov, Tim, et al.
Published: (2025)
Joint Relational Database Generation via Graph-Conditional Diffusion Models
by: Ketata, Mohamed Amine, et al.
Published: (2025)
by: Ketata, Mohamed Amine, et al.
Published: (2025)
Joint Out-of-Distribution Filtering and Data Discovery Active Learning
by: Schmidt, Sebastian, et al.
Published: (2025)
by: Schmidt, Sebastian, et al.
Published: (2025)
Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting
by: Kollovieh, Marcel, et al.
Published: (2024)
by: Kollovieh, Marcel, et al.
Published: (2024)
Uncertainty Estimation for Heterophilic Graphs Through the Lens of Information Theory
by: Fuchsgruber, Dominik, et al.
Published: (2025)
by: Fuchsgruber, Dominik, et al.
Published: (2025)
Effective Data Pruning through Score Extrapolation
by: Schmidt, Sebastian, et al.
Published: (2025)
by: Schmidt, Sebastian, et al.
Published: (2025)
Localized Randomized Smoothing for Collective Robustness Certification
by: Schuchardt, Jan, et al.
Published: (2022)
by: Schuchardt, Jan, et al.
Published: (2022)
Uncertainty for Active Learning on Graphs
by: Fuchsgruber, Dominik, et al.
Published: (2024)
by: Fuchsgruber, Dominik, et al.
Published: (2024)
Expressivity and Generalization: Fragment-Biases for Molecular GNNs
by: Wollschläger, Tom, et al.
Published: (2024)
by: Wollschläger, Tom, et al.
Published: (2024)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026)
by: Saxena, Aman, et al.
Published: (2026)
Unexplored flaws in multiple-choice VQA evaluations
by: Rosenthal, Fabio, et al.
Published: (2025)
by: Rosenthal, Fabio, et al.
Published: (2025)
Randomized Message-Interception Smoothing: Gray-box Certificates for Graph Neural Networks
by: Scholten, Yan, et al.
Published: (2023)
by: Scholten, Yan, et al.
Published: (2023)
Lift Your Molecules: Molecular Graph Generation in Latent Euclidean Space
by: Ketata, Mohamed Amine, et al.
Published: (2024)
by: Ketata, Mohamed Amine, et al.
Published: (2024)
Hierarchical Randomized Smoothing
by: Scholten, Yan, et al.
Published: (2023)
by: Scholten, Yan, et al.
Published: (2023)
Attacking Large Language Models with Projected Gradient Descent
by: Geisler, Simon, et al.
Published: (2024)
by: Geisler, Simon, et al.
Published: (2024)
Certifiably Robust Encoding Schemes
by: Saxena, Aman, et al.
Published: (2024)
by: Saxena, Aman, et al.
Published: (2024)
Discrete Randomized Smoothing Meets Quantum Computing
by: Wollschläger, Tom, et al.
Published: (2024)
by: Wollschläger, Tom, et al.
Published: (2024)
Adversarial Attacks on Graph Neural Networks via Meta Learning
by: Zügner, Daniel, et al.
Published: (2019)
by: Zügner, Daniel, et al.
Published: (2019)
Similar Items
-
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024) -
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
by: Schwinn, Leo, et al.
Published: (2024) -
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025) -
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
by: Scholten, Yan, et al.
Published: (2024) -
Diffusion LLMs are Natural Adversaries for any LLM
by: Lüdke, David, et al.
Published: (2025)