Attacking Large Language Models with Projected Gradient Descent
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Geisler, Simon, Wollschläger, Tom, Abdalla, M. H. I., Gasteiger, Johannes, Günnemann, Stephan |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
par: Geisler, Simon, et autres
Publié: (2025)
par: Geisler, Simon, et autres
Publié: (2025)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
par: Wollschläger, Tom, et autres
Publié: (2025)
par: Wollschläger, Tom, et autres
Publié: (2025)
GemNet: Universal Directional Graph Neural Networks for Molecules
par: Gasteiger, Johannes, et autres
Publié: (2021)
par: Gasteiger, Johannes, et autres
Publié: (2021)
Energy-based Epistemic Uncertainty for Graph Neural Networks
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
What Expressivity Theory Misses: Message Passing Complexity for GNNs
par: Kemper, Niklas, et autres
Publié: (2025)
par: Kemper, Niklas, et autres
Publié: (2025)
Uncertainty Estimation for Heterophilic Graphs Through the Lens of Information Theory
par: Fuchsgruber, Dominik, et autres
Publié: (2025)
par: Fuchsgruber, Dominik, et autres
Publié: (2025)
Graph Neural Networks for Edge Signals: Orientation Equivariance and Invariance
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
Sampling-aware Adversarial Attacks Against Large Language Models
par: Beyer, Tim, et autres
Publié: (2025)
par: Beyer, Tim, et autres
Publié: (2025)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
par: Tomov, Tim, et autres
Publié: (2025)
par: Tomov, Tim, et autres
Publié: (2025)
Spatio-Spectral Graph Neural Networks
par: Geisler, Simon, et autres
Publié: (2024)
par: Geisler, Simon, et autres
Publié: (2024)
Long-Range Graph Wavelet Networks
par: Guerranti, Filippo, et autres
Publié: (2025)
par: Guerranti, Filippo, et autres
Publié: (2025)
SAFT: Structure-Aware Fine-Tuning of LLMs for AMR-to-Text Generation
par: Kamel, Rafiq, et autres
Publié: (2025)
par: Kamel, Rafiq, et autres
Publié: (2025)
Localized Randomized Smoothing for Collective Robustness Certification
par: Schuchardt, Jan, et autres
Publié: (2022)
par: Schuchardt, Jan, et autres
Publié: (2022)
Uncertainty for Active Learning on Graphs
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
par: Fuchsgruber, Dominik, et autres
Publié: (2024)
Expressivity and Generalization: Fragment-Biases for Molecular GNNs
par: Wollschläger, Tom, et autres
Publié: (2024)
par: Wollschläger, Tom, et autres
Publié: (2024)
Diffusion LLMs are Natural Adversaries for any LLM
par: Lüdke, David, et autres
Publié: (2025)
par: Lüdke, David, et autres
Publié: (2025)
Adversarial Robustness of Graph Transformers
par: Foth, Philipp, et autres
Publié: (2024)
par: Foth, Philipp, et autres
Publié: (2024)
Randomized Message-Interception Smoothing: Gray-box Certificates for Graph Neural Networks
par: Scholten, Yan, et autres
Publié: (2023)
par: Scholten, Yan, et autres
Publié: (2023)
Expressivity of Graph Neural Networks Through the Lens of Adversarial Robustness
par: Campi, Francesco, et autres
Publié: (2023)
par: Campi, Francesco, et autres
Publié: (2023)
Lift Your Molecules: Molecular Graph Generation in Latent Euclidean Space
par: Ketata, Mohamed Amine, et autres
Publié: (2024)
par: Ketata, Mohamed Amine, et autres
Publié: (2024)
Certifiably Robust Encoding Schemes
par: Saxena, Aman, et autres
Publié: (2024)
par: Saxena, Aman, et autres
Publié: (2024)
Discrete Randomized Smoothing Meets Quantum Computing
par: Wollschläger, Tom, et autres
Publié: (2024)
par: Wollschläger, Tom, et autres
Publié: (2024)
Explainable Graph Neural Networks Under Fire
par: Li, Zhong, et autres
Publié: (2024)
par: Li, Zhong, et autres
Publié: (2024)
Adversarial Attacks on Graph Neural Networks via Meta Learning
par: Zügner, Daniel, et autres
Publié: (2019)
par: Zügner, Daniel, et autres
Publié: (2019)
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
par: Scholten, Yan, et autres
Publié: (2024)
par: Scholten, Yan, et autres
Publié: (2024)
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
par: Biswas, Sajib, et autres
Publié: (2025)
par: Biswas, Sajib, et autres
Publié: (2025)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
par: Biswas, Sajib, et autres
Publié: (2025)
par: Biswas, Sajib, et autres
Publié: (2025)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
par: Schwinn, Leo, et autres
Publié: (2025)
par: Schwinn, Leo, et autres
Publié: (2025)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
par: Waghela, Hetvi, et autres
Publié: (2024)
par: Waghela, Hetvi, et autres
Publié: (2024)
Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent
par: Wei, Yongxian, et autres
Publié: (2025)
par: Wei, Yongxian, et autres
Publié: (2025)
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
par: Scholten, Yan, et autres
Publié: (2024)
par: Scholten, Yan, et autres
Publié: (2024)
Generative Modeling with Bayesian Sample Inference
par: Lienen, Marten, et autres
Publié: (2025)
par: Lienen, Marten, et autres
Publié: (2025)
Interpolating Discrete Diffusion Models with Controllable Resampling
par: Kollovieh, Marcel, et autres
Publié: (2026)
par: Kollovieh, Marcel, et autres
Publié: (2026)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
par: Gosch, Lukas, et autres
Publié: (2024)
par: Gosch, Lukas, et autres
Publié: (2024)
UnHiPPO: Uncertainty-aware Initialization for State Space Models
par: Lienen, Marten, et autres
Publié: (2025)
par: Lienen, Marten, et autres
Publié: (2025)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
par: Saxena, Aman, et autres
Publié: (2026)
par: Saxena, Aman, et autres
Publié: (2026)
Efficient Adversarial Training in LLMs with Continuous Attacks
par: Xhonneux, Sophie, et autres
Publié: (2024)
par: Xhonneux, Sophie, et autres
Publié: (2024)
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
par: Schwinn, Leo, et autres
Publié: (2024)
par: Schwinn, Leo, et autres
Publié: (2024)
Unfolding Time: Generative Modeling for Turbulent Flows in 4D
par: Saydemir, Abdullah, et autres
Publié: (2024)
par: Saydemir, Abdullah, et autres
Publié: (2024)
On the Inherent Privacy of Zeroth Order Projected Gradient Descent
par: Gupta, Devansh, et autres
Publié: (2025)
par: Gupta, Devansh, et autres
Publié: (2025)
Documents similaires
-
REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective
par: Geisler, Simon, et autres
Publié: (2025) -
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
par: Wollschläger, Tom, et autres
Publié: (2025) -
GemNet: Universal Directional Graph Neural Networks for Molecules
par: Gasteiger, Johannes, et autres
Publié: (2021) -
Energy-based Epistemic Uncertainty for Graph Neural Networks
par: Fuchsgruber, Dominik, et autres
Publié: (2024) -
What Expressivity Theory Misses: Message Passing Complexity for GNNs
par: Kemper, Niklas, et autres
Publié: (2025)