DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection
Fuente:
arXiv
Salvato in:
| Autori principali: | Ren, Junyu, Pan, Xingjian, Gan, Wensheng, Yu, Philip S. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
di: Ying, Zonghao, et al.
Pubblicazione: (2026)
Digital Fingerprinting on Multimedia: A Survey
di: Chen, Wendi, et al.
Pubblicazione: (2024)
di: Chen, Wendi, et al.
Pubblicazione: (2024)
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
di: Wang, Yizhu, et al.
Pubblicazione: (2025)
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026)
di: Cao, Tri, et al.
Pubblicazione: (2026)
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
di: Chen, Sizhe, et al.
Pubblicazione: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
di: Chen, Yulin, et al.
Pubblicazione: (2024)
di: Chen, Yulin, et al.
Pubblicazione: (2024)
StruQ: Defending Against Prompt Injection with Structured Queries
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
PromptSleuth: Detecting Prompt Injection via Semantic Intent Invariance
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
di: Wang, Mengxiao, et al.
Pubblicazione: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
di: Ying, Zonghao, et al.
Pubblicazione: (2025)
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
di: Wu, Xiaodong, et al.
Pubblicazione: (2025)
Strengthening Polymorphic Prompt Assembling: Dynamic Separator Generation Against Emerging Prompt Injection Attacks
di: Dorzhiev, Nima, et al.
Pubblicazione: (2026)
di: Dorzhiev, Nima, et al.
Pubblicazione: (2026)
To Protect the LLM Agent Against the Prompt Injection Attack with Polymorphic Prompt
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
di: Wang, Zhilong, et al.
Pubblicazione: (2025)
Watermarking Techniques for Large Language Models: A Survey
di: Liang, Yuqing, et al.
Pubblicazione: (2024)
di: Liang, Yuqing, et al.
Pubblicazione: (2024)
ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection
di: Weng, Shihao, et al.
Pubblicazione: (2026)
di: Weng, Shihao, et al.
Pubblicazione: (2026)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Securing AI Agents Against Prompt Injection Attacks
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
di: Ramakrishnan, Badrinath, et al.
Pubblicazione: (2025)
Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems
di: Alam, Md Takrim Ul, et al.
Pubblicazione: (2026)
di: Alam, Md Takrim Ul, et al.
Pubblicazione: (2026)
Comparative Analysis Based on DeepSeek, ChatGPT, and Google Gemini: Features, Techniques, Performance, Future Prospects
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
di: Rahman, Anichur, et al.
Pubblicazione: (2025)
The dark deep side of DeepSeek: Fine-tuning attacks against the safety alignment of CoT-enabled models
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
di: Xu, Zhiyuan, et al.
Pubblicazione: (2025)
AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs
di: Wang, Yihan, et al.
Pubblicazione: (2025)
di: Wang, Yihan, et al.
Pubblicazione: (2025)
Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
di: Zhong, Yinan, et al.
Pubblicazione: (2025)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
di: Sehwag, Udari Madhushani, et al.
Pubblicazione: (2026)
di: Sehwag, Udari Madhushani, et al.
Pubblicazione: (2026)
LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents
di: Shah, Harsh
Pubblicazione: (2026)
di: Shah, Harsh
Pubblicazione: (2026)
SecAlign: Defending Against Prompt Injection with Preference Optimization
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
di: Chen, Sizhe, et al.
Pubblicazione: (2024)
Lessons from Defending Gemini Against Indirect Prompt Injections
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
di: Shi, Chongyang, et al.
Pubblicazione: (2025)
Prompt Injection Attack to Tool Selection in LLM Agents
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
di: Shi, Jiawen, et al.
Pubblicazione: (2025)
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
di: Naseh, Ali, et al.
Pubblicazione: (2025)
di: Naseh, Ali, et al.
Pubblicazione: (2025)
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
di: Wang, Jiawen, et al.
Pubblicazione: (2025)
The Vulnerability of LLM Rankers to Prompt Injection Attacks
di: Yin, Yu, et al.
Pubblicazione: (2026)
di: Yin, Yu, et al.
Pubblicazione: (2026)
VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy
di: Cui, Yu, et al.
Pubblicazione: (2025)
di: Cui, Yu, et al.
Pubblicazione: (2025)
A Novel Evaluation Framework for Assessing Resilience Against Prompt Injection Attacks in Large Language Models
di: Yip, Daniel Wankit, et al.
Pubblicazione: (2024)
di: Yip, Daniel Wankit, et al.
Pubblicazione: (2024)
RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
di: Zhong, Peter Yong, et al.
Pubblicazione: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
di: Evtimov, Ivan, et al.
Pubblicazione: (2025)
PROMPTFUZZ: Harnessing Fuzzing Techniques for Robust Testing of Prompt Injection in LLMs
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
di: Yu, Jiahao, et al.
Pubblicazione: (2024)
Node Injection Attack Based on Label Propagation Against Graph Neural Network
di: Zhu, Peican, et al.
Pubblicazione: (2024)
di: Zhu, Peican, et al.
Pubblicazione: (2024)
AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models
di: Wang, Jackson
Pubblicazione: (2026)
di: Wang, Jackson
Pubblicazione: (2026)
PromptShield: Deployable Detection for Prompt Injection Attacks
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
di: Jacob, Dennis, et al.
Pubblicazione: (2025)
Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction
di: Chen, Yulin, et al.
Pubblicazione: (2025)
di: Chen, Yulin, et al.
Pubblicazione: (2025)
ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
di: Lu, Weikai, et al.
Pubblicazione: (2025)
di: Lu, Weikai, et al.
Pubblicazione: (2025)
Signed-Prompt: A New Approach to Prevent Prompt Injection Attacks Against LLM-Integrated Applications
di: Suo, Xuchen
Pubblicazione: (2024)
di: Suo, Xuchen
Pubblicazione: (2024)
Documenti analoghi
-
AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization
di: Ying, Zonghao, et al.
Pubblicazione: (2026) -
Digital Fingerprinting on Multimedia: A Survey
di: Chen, Wendi, et al.
Pubblicazione: (2024) -
Defending Against Prompt Injection with DataFilter
di: Wang, Yizhu, et al.
Pubblicazione: (2025) -
WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections
di: Cao, Tri, et al.
Pubblicazione: (2026) -
Defending Against Prompt Injection With a Few DefensiveTokens
di: Chen, Sizhe, et al.
Pubblicazione: (2025)