Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Fuente:
arXiv
Saved in:
| Main Authors: | Schwinn, Leo, Dobre, David, Xhonneux, Sophie, Gidel, Gauthier, Gunnemann, Stephan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
by: Schwinn, Leo, et al.
Published: (2025)
by: Schwinn, Leo, et al.
Published: (2025)
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025)
by: Scholten, Yan, et al.
Published: (2025)
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025)
by: Dobre, David, et al.
Published: (2025)
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
A Probabilistic Perspective on Unlearning and Alignment for Large Language Models
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
In-Context Learning Can Re-learn Forbidden Tasks
by: Xhonneux, Sophie, et al.
Published: (2024)
by: Xhonneux, Sophie, et al.
Published: (2024)
Extracting Unlearned Information from LLMs with Activation Steering
by: Seyitoğlu, Atakan, et al.
Published: (2024)
by: Seyitoğlu, Atakan, et al.
Published: (2024)
Efficient Time Series Processing for Transformers and State-Space Models through Token Merging
by: Götz, Leon, et al.
Published: (2024)
by: Götz, Leon, et al.
Published: (2024)
Sampling-aware Adversarial Attacks Against Large Language Models
by: Beyer, Tim, et al.
Published: (2025)
by: Beyer, Tim, et al.
Published: (2025)
Diffusion LLMs are Natural Adversaries for any LLM
by: Lüdke, David, et al.
Published: (2025)
by: Lüdke, David, et al.
Published: (2025)
Sarah Frank-Wolfe: Methods for Constrained Optimization with Best Rates and Practical Features
by: Beznosikov, Aleksandr, et al.
Published: (2023)
by: Beznosikov, Aleksandr, et al.
Published: (2023)
Joint Relational Database Generation via Graph-Conditional Diffusion Models
by: Ketata, Mohamed Amine, et al.
Published: (2025)
by: Ketata, Mohamed Amine, et al.
Published: (2025)
Byte Pair Encoding for Efficient Time Series Forecasting
by: Götz, Leon, et al.
Published: (2025)
by: Götz, Leon, et al.
Published: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
Closing the Distribution Gap in Adversarial Training for LLMs
by: Hu, Chengzhi, et al.
Published: (2026)
by: Hu, Chengzhi, et al.
Published: (2026)
Effective Data Pruning through Score Extrapolation
by: Schmidt, Sebastian, et al.
Published: (2025)
by: Schmidt, Sebastian, et al.
Published: (2025)
Joint Out-of-Distribution Filtering and Data Discovery Active Learning
by: Schmidt, Sebastian, et al.
Published: (2025)
by: Schmidt, Sebastian, et al.
Published: (2025)
Flow Matching with Gaussian Process Priors for Probabilistic Time Series Forecasting
by: Kollovieh, Marcel, et al.
Published: (2024)
by: Kollovieh, Marcel, et al.
Published: (2024)
Adversarial Robustness of Graph Transformers
by: Foth, Philipp, et al.
Published: (2024)
by: Foth, Philipp, et al.
Published: (2024)
Unexplored flaws in multiple-choice VQA evaluations
by: Rosenthal, Fabio, et al.
Published: (2025)
by: Rosenthal, Fabio, et al.
Published: (2025)
Assessing Robustness via Score-Based Adversarial Image Generation
by: Kollovieh, Marcel, et al.
Published: (2023)
by: Kollovieh, Marcel, et al.
Published: (2023)
CAP: Controllable Alignment Prompting for Unlearning in LLMs
by: Wang, Zhaokun, et al.
Published: (2026)
by: Wang, Zhaokun, et al.
Published: (2026)
Adversarial Attacks on Graph Neural Networks via Meta Learning
by: Zügner, Daniel, et al.
Published: (2019)
by: Zügner, Daniel, et al.
Published: (2019)
A Scalable Multi-Task Model for Virtual Sensors
by: Götz, Leon, et al.
Published: (2026)
by: Götz, Leon, et al.
Published: (2026)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Advantage Alignment Algorithms
by: Duque, Juan Agustin, et al.
Published: (2024)
by: Duque, Juan Agustin, et al.
Published: (2024)
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
by: Spohn, Philipp, et al.
Published: (2025)
by: Spohn, Philipp, et al.
Published: (2025)
UnHiPPO: Uncertainty-aware Initialization for State Space Models
by: Lienen, Marten, et al.
Published: (2025)
by: Lienen, Marten, et al.
Published: (2025)
Bypassing the Safety Training of Open-Source LLMs with Priming Attacks
by: Vega, Jason, et al.
Published: (2023)
by: Vega, Jason, et al.
Published: (2023)
Large-Scale Dataset Pruning in Adversarial Training through Data Importance Extrapolation
by: Nieth, Björn, et al.
Published: (2024)
by: Nieth, Björn, et al.
Published: (2024)
Performative Prediction with Neural Networks
by: Mofakhami, Mehrnaz, et al.
Published: (2023)
by: Mofakhami, Mehrnaz, et al.
Published: (2023)
Proving Linear Mode Connectivity of Neural Networks via Optimal Transport
by: Ferbach, Damien, et al.
Published: (2023)
by: Ferbach, Damien, et al.
Published: (2023)
Unlocking Point Processes through Point Set Diffusion
by: Lüdke, David, et al.
Published: (2024)
by: Lüdke, David, et al.
Published: (2024)
Provably Reliable Conformal Prediction Sets in the Presence of Data Poisoning
by: Scholten, Yan, et al.
Published: (2024)
by: Scholten, Yan, et al.
Published: (2024)
Discrete Compositional Generation via General Soft Operators and Robust Reinforcement Learning
by: Jiralerspong, Marco, et al.
Published: (2025)
by: Jiralerspong, Marco, et al.
Published: (2025)
Omega: Optimistic EMA Gradients
by: Ramirez, Juan, et al.
Published: (2023)
by: Ramirez, Juan, et al.
Published: (2023)
Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
by: Gosch, Lukas, et al.
Published: (2024)
by: Gosch, Lukas, et al.
Published: (2024)
Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy
by: Saxena, Aman, et al.
Published: (2026)
by: Saxena, Aman, et al.
Published: (2026)
Transferable SCF-Acceleration through Solver-Aligned Initialization Learning
by: Eberhard, Eike S., et al.
Published: (2026)
by: Eberhard, Eike S., et al.
Published: (2026)
Similar Items
-
Efficient Adversarial Training in LLMs with Continuous Attacks
by: Xhonneux, Sophie, et al.
Published: (2024) -
Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives
by: Schwinn, Leo, et al.
Published: (2025) -
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
by: Scholten, Yan, et al.
Published: (2025) -
A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens
by: Dobre, David, et al.
Published: (2025) -
LLM-Safety Evaluations Lack Robustness
by: Beyer, Tim, et al.
Published: (2025)