Counterfactual Explainable Incremental Prompt Attack Analysis on Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Shu, Dong, Jin, Mingyu, Chen, Tianle, Zhang, Chong, Zhang, Yongfeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024)
by: Zhang, Chong, et al.
Published: (2024)
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
by: Xiong, Junjie, et al.
Published: (2025)
by: Xiong, Junjie, et al.
Published: (2025)
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024)
by: Sha, Zeyang, et al.
Published: (2024)
Prompt Inversion Attack against Collaborative Inference of Large Language Models
by: Qu, Wenjie, et al.
Published: (2025)
by: Qu, Wenjie, et al.
Published: (2025)
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
by: Peng, Benji, et al.
Published: (2024)
by: Peng, Benji, et al.
Published: (2024)
System Prompt Extraction Attacks and Defenses in Large Language Models
by: Das, Badhan Chandra, et al.
Published: (2025)
by: Das, Badhan Chandra, et al.
Published: (2025)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
by: Zhao, Shiqian, et al.
Published: (2025)
by: Zhao, Shiqian, et al.
Published: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Prompt Inference Attack on Distributed Large Language Model Inference Frameworks
by: Luo, Xinjian, et al.
Published: (2025)
by: Luo, Xinjian, et al.
Published: (2025)
Learnable Linguistic Watermarks for Tracing Model Extraction Attacks on Large Language Models
by: Bai, Minhao, et al.
Published: (2024)
by: Bai, Minhao, et al.
Published: (2024)
PRIVMARK: Private Large Language Models Watermarking with MPC
by: Fargues, Thomas, et al.
Published: (2025)
by: Fargues, Thomas, et al.
Published: (2025)
Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models
by: Chen, Zeyuan, et al.
Published: (2026)
by: Chen, Zeyuan, et al.
Published: (2026)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
by: Wang, Jackson
Published: (2026)
by: Wang, Jackson
Published: (2026)
Detecting Complex Multi-step Attacks with Explainable Graph Neural Network
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
MorphMark: Flexible Adaptive Watermarking for Large Language Models
by: Wang, Zongqi, et al.
Published: (2025)
by: Wang, Zongqi, et al.
Published: (2025)
Assessing Prompt Injection Risks in 200+ Custom GPTs
by: Yu, Jiahao, et al.
Published: (2023)
by: Yu, Jiahao, et al.
Published: (2023)
An Engorgio Prompt Makes Large Language Model Babble on
by: Dong, Jianshuo, et al.
Published: (2024)
by: Dong, Jianshuo, et al.
Published: (2024)
Safeguarding Large Language Models: A Survey
by: Dong, Yi, et al.
Published: (2024)
by: Dong, Yi, et al.
Published: (2024)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks
by: Khouna, Awa, et al.
Published: (2025)
by: Khouna, Awa, et al.
Published: (2025)
A Novel Evaluation Framework for Assessing Resilience Against Prompt Injection Attacks in Large Language Models
by: Yip, Daniel Wankit, et al.
Published: (2024)
by: Yip, Daniel Wankit, et al.
Published: (2024)
AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2026)
by: Jia, Yuqi, et al.
Published: (2026)
Explainable and Transferable Adversarial Attack for ML-Based Network Intrusion Detectors
by: Zhang, Hangsheng, et al.
Published: (2024)
by: Zhang, Hangsheng, et al.
Published: (2024)
SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
by: Collu, Matteo Gioele, et al.
Published: (2025)
by: Collu, Matteo Gioele, et al.
Published: (2025)
JBShield: Defending Large Language Models from Jailbreak Attacks through Activated Concept Analysis and Manipulation
by: Zhang, Shenyi, et al.
Published: (2025)
by: Zhang, Shenyi, et al.
Published: (2025)
SOS! Soft Prompt Attack Against Open-Source Large Language Models
by: Yang, Ziqing, et al.
Published: (2024)
by: Yang, Ziqing, et al.
Published: (2024)
Membership Inference Attacks on Tokenizers of Large Language Models
by: Tong, Meng, et al.
Published: (2025)
by: Tong, Meng, et al.
Published: (2025)
EmojiPrompt: Generative Prompt Obfuscation for Privacy-Preserving Communication with Cloud-based LLMs
by: Lin, Sam, et al.
Published: (2024)
by: Lin, Sam, et al.
Published: (2024)
Shifting-Merging: Secure, High-Capacity and Efficient Steganography via Large Language Models
by: Bai, Minhao, et al.
Published: (2025)
by: Bai, Minhao, et al.
Published: (2025)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
by: Pang, Kaiyi, et al.
Published: (2024)
by: Pang, Kaiyi, et al.
Published: (2024)
Automating Prompt Leakage Attacks on Large Language Models Using Agentic Approach
by: Sternak, Tvrtko, et al.
Published: (2025)
by: Sternak, Tvrtko, et al.
Published: (2025)
Prompt Injection Attacks on Large Language Models in Oncology
by: Clusmann, Jan, et al.
Published: (2024)
by: Clusmann, Jan, et al.
Published: (2024)
Token Highlighter: Inspecting and Mitigating Jailbreak Prompts for Large Language Models
by: Hu, Xiaomeng, et al.
Published: (2024)
by: Hu, Xiaomeng, et al.
Published: (2024)
PINA: Prompt Injection Attack against Navigation Agents
by: Liu, Jiani, et al.
Published: (2026)
by: Liu, Jiani, et al.
Published: (2026)
Semantic Steganography: A Framework for Robust and High-Capacity Information Hiding using Large Language Models
by: Bai, Minhao, et al.
Published: (2024)
by: Bai, Minhao, et al.
Published: (2024)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Risk Assessment and Security Analysis of Large Language Models
by: Zhang, Xiaoyan, et al.
Published: (2025)
by: Zhang, Xiaoyan, et al.
Published: (2025)
Towards Robust Multimodal Large Language Models Against Jailbreak Attacks
by: Yin, Ziyi, et al.
Published: (2025)
by: Yin, Ziyi, et al.
Published: (2025)
CPA-RAG:Covert Poisoning Attacks on Retrieval-Augmented Generation in Large Language Models
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
Similar Items
-
Goal-guided Generative Prompt Injection Attack on Large Language Models
by: Zhang, Chong, et al.
Published: (2024) -
Invisible Prompts, Visible Threats: Malicious Font Injection in External Resources for Large Language Models
by: Xiong, Junjie, et al.
Published: (2025) -
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024) -
Prompt Inversion Attack against Collaborative Inference of Large Language Models
by: Qu, Wenjie, et al.
Published: (2025) -
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
by: Peng, Benji, et al.
Published: (2024)