Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Biswas, Sajib, Nishino, Mao, Chacko, Samuel Jacob, Liu, Xiuwen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
von: Biswas, Sajib, et al.
Veröffentlicht: (2025)
Adversarial Attacks on Large Language Models Using Regularized Relaxation
von: Chacko, Samuel Jacob, et al.
Veröffentlicht: (2024)
von: Chacko, Samuel Jacob, et al.
Veröffentlicht: (2024)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
von: Tan, Yuting, et al.
Veröffentlicht: (2025)
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025)
von: Li, Ran, et al.
Veröffentlicht: (2025)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
von: Roth, Tom, et al.
Veröffentlicht: (2021)
von: Roth, Tom, et al.
Veröffentlicht: (2021)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
User Inference Attacks on Large Language Models
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
von: Kandpal, Nikhil, et al.
Veröffentlicht: (2023)
Composite Backdoor Attacks Against Large Language Models
von: Huang, Hai, et al.
Veröffentlicht: (2023)
von: Huang, Hai, et al.
Veröffentlicht: (2023)
Hijacking Large Language Models via Adversarial In-Context Learning
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
von: Zhou, Xiangyu, et al.
Veröffentlicht: (2023)
TFL: Targeted Bit-Flip Attack on Large Language Model
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
von: Guo, Jingkai, et al.
Veröffentlicht: (2026)
An Early Categorization of Prompt Injection Attacks on Large Language Models
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
von: Rossi, Sippo, et al.
Veröffentlicht: (2024)
MARAGE: Transferable Multi-Model Adversarial Attack for Retrieval-Augmented Generation Data Extraction
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
von: Hu, Xiao, et al.
Veröffentlicht: (2025)
Confidence Elicitation: A New Attack Vector for Large Language Models
von: Formento, Brian, et al.
Veröffentlicht: (2025)
von: Formento, Brian, et al.
Veröffentlicht: (2025)
Securing Large Language Models (LLMs) from Prompt Injection Attacks
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
von: Suri, Omar Farooq Khan, et al.
Veröffentlicht: (2025)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
von: Zhu, Kaijie, et al.
Veröffentlicht: (2023)
SBFA: Single Sneaky Bit Flip Attack to Break Large Language Models
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
von: Guo, Jingkai, et al.
Veröffentlicht: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
von: Yang, Ziqing, et al.
Veröffentlicht: (2024)
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
von: Hu, Xiaomeng, et al.
Veröffentlicht: (2024)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024)
Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
von: Fu, Wenjie, et al.
Veröffentlicht: (2023)
Cyber-Attack Technique Classification Using Two-Stage Trained Large Language Models
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
von: You, Weiqiu, et al.
Veröffentlicht: (2024)
IDT: Dual-Task Adversarial Attacks for Privacy Protection
von: Faustini, Pedro, et al.
Veröffentlicht: (2024)
von: Faustini, Pedro, et al.
Veröffentlicht: (2024)
Towards More Realistic Extraction Attacks: An Adversarial Perspective
von: More, Yash, et al.
Veröffentlicht: (2024)
von: More, Yash, et al.
Veröffentlicht: (2024)
Cross-Entropy Attacks to Language Models via Rare Event Simulation
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
von: Ni, Mingze, et al.
Veröffentlicht: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
von: Brown, Hannah, et al.
Veröffentlicht: (2024)
Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models
von: He, Jiaming, et al.
Veröffentlicht: (2024)
von: He, Jiaming, et al.
Veröffentlicht: (2024)
Adversarial Vulnerabilities in Large Language Models for Time Series Forecasting
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
von: Liu, Fuqiang, et al.
Veröffentlicht: (2024)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
von: Dey, Roopkatha, et al.
Veröffentlicht: (2024)
Deciphering the Chaos: Enhancing Jailbreak Attacks via Adversarial Prompt Translation
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
von: Li, Qizhang, et al.
Veröffentlicht: (2024)
Beyond Gradient and Priors in Privacy Attacks: Leveraging Pooler Layer Inputs of Language Models in Federated Learning
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
Adversarial Text Purification: A Large Language Model Approach for Defense
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
von: Moraffah, Raha, et al.
Veröffentlicht: (2024)
MPAT: Building Robust Deep Neural Networks against Textual Adversarial Attacks
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Fangyuan, et al.
Veröffentlicht: (2024)
Differentially Private Tabular Data Synthesis using Large Language Models
von: Tran, Toan V., et al.
Veröffentlicht: (2024)
von: Tran, Toan V., et al.
Veröffentlicht: (2024)
Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2023)
Humanizing Machine-Generated Content: Evading AI-Text Detection through Adversarial Attack
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
von: Zhou, Ying, et al.
Veröffentlicht: (2024)
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
von: Yi, Sibo, et al.
Veröffentlicht: (2024)
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
von: Xue, Eric, et al.
Veröffentlicht: (2025)
von: Xue, Eric, et al.
Veröffentlicht: (2025)
Exploring Vulnerabilities and Protections in Large Language Models: A Survey
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
von: Liu, Frank Weizhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
von: Biswas, Sajib, et al.
Veröffentlicht: (2025) -
Adversarial Attacks on Large Language Models Using Regularized Relaxation
von: Chacko, Samuel Jacob, et al.
Veröffentlicht: (2024) -
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
von: Waghela, Hetvi, et al.
Veröffentlicht: (2024) -
The Resurgence of GCG Adversarial Attacks on Large Language Models
von: Tan, Yuting, et al.
Veröffentlicht: (2025) -
LARGO: Latent Adversarial Reflection through Gradient Optimization for Jailbreaking LLMs
von: Li, Ran, et al.
Veröffentlicht: (2025)