Sampling-aware Adversarial Attacks Against Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Beyer, Tim, Scholten, Yan, Schwinn, Leo, Günnemann, Stephan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
von: Scholten, Yan, et al.
Veröffentlicht: (2025)
von: Scholten, Yan, et al.
Veröffentlicht: (2025)
Efficient Adversarial Training in LLMs with Continuous Attacks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024)
Assessing Robustness via Score-Based Adversarial Image Generation
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2023)
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2023)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
von: Schwinn, Leo, et al.
Veröffentlicht: (2025)
von: Schwinn, Leo, et al.
Veröffentlicht: (2025)
Adversarial Robustness of Graph Transformers
von: Foth, Philipp, et al.
Veröffentlicht: (2024)
von: Foth, Philipp, et al.
Veröffentlicht: (2024)
Diffusion LLMs are Natural Adversaries for any LLM
von: Lüdke, David, et al.
Veröffentlicht: (2025)
von: Lüdke, David, et al.
Veröffentlicht: (2025)
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
von: Scholten, Yan, et al.
Veröffentlicht: (2024)
Efficient Time Series Processing for Transformers and State-Space Models through Token Merging
von: Götz, Leon, et al.
Veröffentlicht: (2024)
von: Götz, Leon, et al.
Veröffentlicht: (2024)
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More
von: Schuchardt, Jan, et al.
Veröffentlicht: (2023)
von: Schuchardt, Jan, et al.
Veröffentlicht: (2023)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
von: Schwinn, Leo, et al.
Veröffentlicht: (2024)
von: Schwinn, Leo, et al.
Veröffentlicht: (2024)
Joint Relational Database Generation via Graph-Conditional Diffusion Models
von: Ketata, Mohamed Amine, et al.
Veröffentlicht: (2025)
von: Ketata, Mohamed Amine, et al.
Veröffentlicht: (2025)
Byte Pair Encoding for Efficient Time Series Forecasting
von: Götz, Leon, et al.
Veröffentlicht: (2025)
von: Götz, Leon, et al.
Veröffentlicht: (2025)
Closing the Distribution Gap in Adversarial Training for LLMs
von: Hu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Hu, Chengzhi, et al.
Veröffentlicht: (2026)
Joint Out-of-Distribution Filtering and Data Discovery Active Learning
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
Extracting Unlearned Information from LLMs with Activation Steering
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
von: Seyitoğlu, Atakan, et al.
Veröffentlicht: (2024)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
von: Saxena, Aman, et al.
Veröffentlicht: (2026)
von: Saxena, Aman, et al.
Veröffentlicht: (2026)
Adversarial Attacks on Graph Neural Networks via Meta Learning
von: Zügner, Daniel, et al.
Veröffentlicht: (2019)
von: Zügner, Daniel, et al.
Veröffentlicht: (2019)
Fast Proxies for LLM Robustness Evaluation
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2024)
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2024)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
von: Campi, Francesco, et al.
Veröffentlicht: (2023)
von: Campi, Francesco, et al.
Veröffentlicht: (2023)
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
von: Geisler, Simon, et al.
Veröffentlicht: (2025)
von: Geisler, Simon, et al.
Veröffentlicht: (2025)
Effective Data Pruning through Score Extrapolation
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
Attacking Large Language Models with Projected Gradient Descent
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
von: Geisler, Simon, et al.
Veröffentlicht: (2024)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
von: Schwinn, Leo, et al.
Veröffentlicht: (2026)
Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
von: Nieth, Björn, et al.
Veröffentlicht: (2024)
von: Nieth, Björn, et al.
Veröffentlicht: (2024)
A Scalable Multi-Task Model for Virtual Sensors
von: Götz, Leon, et al.
Veröffentlicht: (2026)
von: Götz, Leon, et al.
Veröffentlicht: (2026)
Unexplored flaws in multiple-choice VQA evaluations
von: Rosenthal, Fabio, et al.
Veröffentlicht: (2025)
von: Rosenthal, Fabio, et al.
Veröffentlicht: (2025)
UnHiPPO: Uncertainty-aware Initialization for State Space Models
von: Lienen, Marten, et al.
Veröffentlicht: (2025)
von: Lienen, Marten, et al.
Veröffentlicht: (2025)
Randomized Message-Interception Smoothing: Gray-box Certificates for Graph Neural Networks
von: Scholten, Yan, et al.
Veröffentlicht: (2023)
von: Scholten, Yan, et al.
Veröffentlicht: (2023)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
von: Gosch, Lukas, et al.
Veröffentlicht: (2024)
von: Gosch, Lukas, et al.
Veröffentlicht: (2024)
Generative Modeling with Bayesian Sample Inference
von: Lienen, Marten, et al.
Veröffentlicht: (2025)
von: Lienen, Marten, et al.
Veröffentlicht: (2025)
Hierarchical Randomized Smoothing
von: Scholten, Yan, et al.
Veröffentlicht: (2023)
von: Scholten, Yan, et al.
Veröffentlicht: (2023)
Task-Awareness Improves LLM Generations and Uncertainty
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
LLM-Safety Evaluations Lack Robustness
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
AdversariaLLM: A Unified and Modular Toolbox for LLM Robustness Research
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
von: Beyer, Tim, et al.
Veröffentlicht: (2025)
Discrete Bayesian Sample Inference for Graph Generation
von: Petersen, Ole, et al.
Veröffentlicht: (2025)
von: Petersen, Ole, et al.
Veröffentlicht: (2025)
Exact Certification of (Graph) Neural Networks Against Label Poisoning
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
von: Sabanayagam, Mahalakshmi, et al.
Veröffentlicht: (2024)
Graph Neural Networks for Edge Signals: Orientation Equivariance and Invariance
von: Fuchsgruber, Dominik, et al.
Veröffentlicht: (2024)
von: Fuchsgruber, Dominik, et al.
Veröffentlicht: (2024)
Interpolating Discrete Diffusion Models with Controllable Resampling
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2026)
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
von: Scholten, Yan, et al.
Veröffentlicht: (2024) -
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
von: Scholten, Yan, et al.
Veröffentlicht: (2025) -
Efficient Adversarial Training in LLMs with Continuous Attacks
von: Xhonneux, Sophie, et al.
Veröffentlicht: (2024) -
Assessing Robustness via Score-Based Adversarial Image Generation
von: Kollovieh, Marcel, et al.
Veröffentlicht: (2023) -
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
von: Schwinn, Leo, et al.
Veröffentlicht: (2025)