Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Xiao, Li, Liangzhi, Xiang, Tong, Ye, Fuying, Wei, Lu, Li, Wangyue, Garcia, Noa |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
by: Ni, Zhenyang, et al.
Published: (2024)
by: Ni, Zhenyang, et al.
Published: (2024)
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
by: Wang, Nan, et al.
Published: (2025)
by: Wang, Nan, et al.
Published: (2025)
Exploiting AI for Attacks: On the Interplay between Adversarial AI and Offensive AI
by: Schröer, Saskia Laura, et al.
Published: (2025)
by: Schröer, Saskia Laura, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
by: Liu, Hongfu, et al.
Published: (2024)
by: Liu, Hongfu, et al.
Published: (2024)
ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected
by: Gharami, Kanchon, et al.
Published: (2025)
by: Gharami, Kanchon, et al.
Published: (2025)
QueryAttack: Jailbreaking Aligned Large Language Models Using Structured Non-natural Query Language
by: Zou, Qingsong, et al.
Published: (2025)
by: Zou, Qingsong, et al.
Published: (2025)
Privacy Leaks by Adversaries: Adversarial Iterations for Membership Inference Attack
by: Xue, Jing, et al.
Published: (2025)
by: Xue, Jing, et al.
Published: (2025)
Membership Inference Attacks Against Video Large Language Models
by: Song, Wei, et al.
Published: (2026)
by: Song, Wei, et al.
Published: (2026)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2026)
by: Jia, Yuqi, et al.
Published: (2026)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks
by: Tong, Haibo, et al.
Published: (2025)
by: Tong, Haibo, et al.
Published: (2025)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
NatGVD: Natural Adversarial Example Attack towards Graph-based Vulnerability Detection
by: Rath, Avilash, et al.
Published: (2025)
by: Rath, Avilash, et al.
Published: (2025)
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
by: Collu, Matteo Gioele, et al.
Published: (2025)
by: Collu, Matteo Gioele, et al.
Published: (2025)
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
by: Xiang, Kunlan, et al.
Published: (2025)
by: Xiang, Kunlan, et al.
Published: (2025)
Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models
by: Li, Xiao, et al.
Published: (2024)
by: Li, Xiao, et al.
Published: (2024)
Tit-for-Tat: Safeguarding Large Vision-Language Models Against Jailbreak Attacks via Adversarial Defense
by: Hao, Shuyang, et al.
Published: (2025)
by: Hao, Shuyang, et al.
Published: (2025)
Quantization Aware Attack: Enhancing Transferable Adversarial Attacks by Model Quantization
by: Yang, Yulong, et al.
Published: (2023)
by: Yang, Yulong, et al.
Published: (2023)
Prompt Inference Attack on Distributed Large Language Model Inference Frameworks
by: Luo, Xinjian, et al.
Published: (2025)
by: Luo, Xinjian, et al.
Published: (2025)
The Resurgence of GCG Adversarial Attacks on Large Language Models
by: Tan, Yuting, et al.
Published: (2025)
by: Tan, Yuting, et al.
Published: (2025)
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
by: Zhao, Zhengyue, et al.
Published: (2025)
by: Zhao, Zhengyue, et al.
Published: (2025)
Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey
by: Jain, Bhavuk, et al.
Published: (2026)
by: Jain, Bhavuk, et al.
Published: (2026)
Robustness Against Adversarial Attacks via Learning Confined Adversarial Polytopes
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
by: Hamidi, Shayan Mohajer, et al.
Published: (2024)
Double Backdoored: Converting Code Large Language Model Backdoors to Traditional Malware via Adversarial Instruction Tuning Attacks
by: Hossen, Md Imran, et al.
Published: (2024)
by: Hossen, Md Imran, et al.
Published: (2024)
Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
An Automated Attack Investigation Approach Leveraging Threat-Knowledge-Augmented Large Language Models
by: Dai, Rujie, et al.
Published: (2025)
by: Dai, Rujie, et al.
Published: (2025)
External Data Extraction Attacks against Retrieval-Augmented Large Language Models
by: He, Yu, et al.
Published: (2025)
by: He, Yu, et al.
Published: (2025)
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
ALRPHFS: Adversarially Learned Risk Patterns with Hierarchical Fast \& Slow Reasoning for Robust Agent Defense
by: Xiang, Shiyu, et al.
Published: (2025)
by: Xiang, Shiyu, et al.
Published: (2025)
Can Reinforcement Learning Unlock the Hidden Dangers in Aligned Large Language Models?
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
by: Karkevandi, Mohammad Bahrami, et al.
Published: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Invisible Adversaries: A Systematic Study of Session Manipulation Attacks on VPNs
by: Yang, Yuxiang, et al.
Published: (2026)
by: Yang, Yuxiang, et al.
Published: (2026)
Safety Layers in Aligned Large Language Models: The Key to LLM Security
by: Li, Shen, et al.
Published: (2024)
by: Li, Shen, et al.
Published: (2024)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
by: Peng, Benji, et al.
Published: (2024)
by: Peng, Benji, et al.
Published: (2024)
Mitigating Adversarial Effects of False Data Injection Attacks in Power Grid
by: Riya, Farhin Farhad, et al.
Published: (2023)
by: Riya, Farhin Farhad, et al.
Published: (2023)
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial Purification
by: Kang, Mintong, et al.
Published: (2023)
by: Kang, Mintong, et al.
Published: (2023)
RHINO: Guided Reasoning for Mapping Network Logs to Adversarial Tactics and Techniques with Large Language Models
by: Meng, Fanchao, et al.
Published: (2025)
by: Meng, Fanchao, et al.
Published: (2025)
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
by: Jiang, Hanrui, et al.
Published: (2025)
by: Jiang, Hanrui, et al.
Published: (2025)
Similar Items
-
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language Models
by: Ni, Zhenyang, et al.
Published: (2024) -
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
by: Wang, Nan, et al.
Published: (2025) -
Exploiting AI for Attacks: On the Interplay between Adversarial AI and Offensive AI
by: Schröer, Saskia Laura, et al.
Published: (2025) -
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024) -
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
by: Liu, Hongfu, et al.
Published: (2024)