Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hongfu, Xie, Yuxi, Wang, Ye, Shieh, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
von: Tang, Haochun, et al.
Veröffentlicht: (2026)
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
Learning to Poison Large Language Models for Downstream Manipulation
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
Privacy Preserving In-Context-Learning Framework for Large Language Models
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
von: Bhusal, Bishnu, et al.
Veröffentlicht: (2025)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Digger: Detecting Copyright Content Mis-usage in Large Language Model Training
von: Li, Haodong, et al.
Veröffentlicht: (2024)
von: Li, Haodong, et al.
Veröffentlicht: (2024)
Token-Specific Watermarking with Enhanced Detectability and Semantic Coherence for Large Language Models
von: Huo, Mingjia, et al.
Veröffentlicht: (2024)
von: Huo, Mingjia, et al.
Veröffentlicht: (2024)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
CR-UTP: Certified Robustness against Universal Text Perturbations on Large Language Models
von: Lou, Qian, et al.
Veröffentlicht: (2024)
von: Lou, Qian, et al.
Veröffentlicht: (2024)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
von: Roth, Tom, et al.
Veröffentlicht: (2021)
von: Roth, Tom, et al.
Veröffentlicht: (2021)
MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
von: Fu, Wenjie, et al.
Veröffentlicht: (2024)
Privately Learning from Graphs with Applications in Fine-tuning Large Language Models
von: Yin, Haoteng, et al.
Veröffentlicht: (2024)
von: Yin, Haoteng, et al.
Veröffentlicht: (2024)
Network Traffic Classification Using Machine Learning, Transformer, and Large Language Models
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
von: Antari, Ahmad, et al.
Veröffentlicht: (2025)
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
von: Fan, Mingyuan, et al.
Veröffentlicht: (2026)
Detecting Pretraining Data from Large Language Models
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
von: Shi, Weijia, et al.
Veröffentlicht: (2023)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Model Provenance Testing for Large Language Models
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
von: Nikolic, Ivica, et al.
Veröffentlicht: (2025)
A Watermark for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
On the Reliability of Watermarks for Large Language Models
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
von: Kirchenbauer, John, et al.
Veröffentlicht: (2023)
Improved Techniques for Optimization-Based Jailbreaking on Large Language Models
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
von: Jia, Xiaojun, et al.
Veröffentlicht: (2024)
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2025)
Malware Classification from Memory Dumps Using Machine Learning, Transformers, and Large Language Models
von: Dweib, Areej, et al.
Veröffentlicht: (2025)
von: Dweib, Areej, et al.
Veröffentlicht: (2025)
Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models
von: Fu, Wenjie, et al.
Veröffentlicht: (2025)
von: Fu, Wenjie, et al.
Veröffentlicht: (2025)
RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models
von: Chugh, Rishit
Veröffentlicht: (2026)
von: Chugh, Rishit
Veröffentlicht: (2026)
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
von: Khachaturov, David, et al.
Veröffentlicht: (2025)
Topic-Based Watermarks for Large Language Models
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
von: Nemecek, Alexander, et al.
Veröffentlicht: (2024)
Representation Bending for Large Language Model Safety
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
von: Yousefpour, Ashkan, et al.
Veröffentlicht: (2025)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Debiasing Watermarks for Large Language Models via Maximal Coupling
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2024)
von: Xie, Yangxinyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024) -
On Calibration of LLM-based Guard Models for Reliable Content Moderation
von: Liu, Hongfu, et al.
Veröffentlicht: (2024) -
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023) -
Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
von: Tang, Haochun, et al.
Veröffentlicht: (2026) -
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)