Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Sahu, Anubhab, Samanta, Diptisha, Soosahabi, Reza |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
by: Zhao, Jin, et al.
Published: (2026)
by: Zhao, Jin, et al.
Published: (2026)
DeepNcode: Encoding-Based Protection against Bit-Flip Attacks on Neural Networks
by: Velčický, Patrik, et al.
Published: (2024)
by: Velčický, Patrik, et al.
Published: (2024)
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026)
by: Ou, Haoran, et al.
Published: (2026)
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025)
by: Li, Yige, et al.
Published: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)
by: Jia, Yuqi, et al.
Published: (2025)
RvB: Automating AI System Hardening via Iterative Red-Blue Games
by: Huang, Lige, et al.
Published: (2026)
by: Huang, Lige, et al.
Published: (2026)
Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems
by: Jiao, Ruochen, et al.
Published: (2024)
by: Jiao, Ruochen, et al.
Published: (2024)
UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks
by: Yu, Tianlong, et al.
Published: (2026)
by: Yu, Tianlong, et al.
Published: (2026)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
by: Xu, Yixiao, et al.
Published: (2025)
by: Xu, Yixiao, et al.
Published: (2025)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
by: Zhang, Kaiyuan, et al.
Published: (2025)
by: Zhang, Kaiyuan, et al.
Published: (2025)
Cascade: Composing Software-Hardware Attack Gadgets for Adversarial Threat Amplification in Compound AI Systems
by: Banerjee, Sarbartha, et al.
Published: (2026)
by: Banerjee, Sarbartha, et al.
Published: (2026)
CheatAgent: Attacking LLM-Empowered Recommender Systems via LLM Agent
by: Ning, Liang-bo, et al.
Published: (2025)
by: Ning, Liang-bo, et al.
Published: (2025)
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
AutoPentester: An LLM Agent-based Framework for Automated Pentesting
by: Ginige, Yasod, et al.
Published: (2025)
by: Ginige, Yasod, et al.
Published: (2025)
Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model
by: Wang, Tianyi, et al.
Published: (2026)
by: Wang, Tianyi, et al.
Published: (2026)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems
by: Balashov, Andrii, et al.
Published: (2025)
by: Balashov, Andrii, et al.
Published: (2025)
Model Inversion Attack against Federated Unlearning
by: Zhou, Lei, et al.
Published: (2025)
by: Zhou, Lei, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
by: Liu, Shuyuan, et al.
Published: (2025)
by: Liu, Shuyuan, et al.
Published: (2025)
Semantic-Aware Advanced Persistent Threat Detection Using Autoencoders on LLM-Encoded System Logs
by: Mohammed, Waleed Khan, et al.
Published: (2026)
by: Mohammed, Waleed Khan, et al.
Published: (2026)
Defending against Indirect Prompt Injection by Instruction Detection
by: Wen, Tongyu, et al.
Published: (2025)
by: Wen, Tongyu, et al.
Published: (2025)
E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems
by: Guan, Zelin, et al.
Published: (2026)
by: Guan, Zelin, et al.
Published: (2026)
SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
by: Sivaroopan, Nirhoshan, et al.
Published: (2026)
RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework
by: Ikbarieh, Seif, et al.
Published: (2025)
by: Ikbarieh, Seif, et al.
Published: (2025)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
by: Truong, Vu Tuan, et al.
Published: (2026)
by: Truong, Vu Tuan, et al.
Published: (2026)
Optimizing Adaptive Attacks against Watermarks for Language Models
by: Diaa, Abdulrahman, et al.
Published: (2024)
by: Diaa, Abdulrahman, et al.
Published: (2024)
TEAM: Temporal Adversarial Examples Attack Model against Network Intrusion Detection System Applied to RNN
by: Liu, Ziyi, et al.
Published: (2024)
by: Liu, Ziyi, et al.
Published: (2024)
QueryCheetah: Fast Automated Discovery of Attribute Inference Attacks Against Query-Based Systems
by: Stevanoski, Bozhidar, et al.
Published: (2024)
by: Stevanoski, Bozhidar, et al.
Published: (2024)
LLM-based Multi-class Attack Analysis and Mitigation Framework in IoT/IIoT Networks
by: Ikbarieh, Seif, et al.
Published: (2025)
by: Ikbarieh, Seif, et al.
Published: (2025)
How to Train your Antivirus: RL-based Hardening through the Problem-Space
by: Tsingenopoulos, Ilias, et al.
Published: (2024)
by: Tsingenopoulos, Ilias, et al.
Published: (2024)
SAGE: A Generic Framework for LLM Safety Evaluation
by: Jindal, Madhur, et al.
Published: (2025)
by: Jindal, Madhur, et al.
Published: (2025)
When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents
by: Liu, Shi, et al.
Published: (2026)
by: Liu, Shi, et al.
Published: (2026)
WeiDetect: Weibull Distribution-Based Defense against Poisoning Attacks in Federated Learning for Network Intrusion Detection Systems
by: M., Sameera K., et al.
Published: (2025)
by: M., Sameera K., et al.
Published: (2025)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Semantic-level Backdoor Attack against Text-to-Image Diffusion Models
by: Chen, Tianxin, et al.
Published: (2026)
by: Chen, Tianxin, et al.
Published: (2026)
FedCC: Robust Federated Learning against Model Poisoning Attacks
by: Jeong, Hyejun, et al.
Published: (2022)
by: Jeong, Hyejun, et al.
Published: (2022)
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
by: Li, Yuying, et al.
Published: (2024)
by: Li, Yuying, et al.
Published: (2024)
CUBA: Controlled Untargeted Backdoor Attack against Deep Neural Networks
by: Wu, Yinghao, et al.
Published: (2025)
by: Wu, Yinghao, et al.
Published: (2025)
Similar Items
-
Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks
by: Zhao, Jin, et al.
Published: (2026) -
DeepNcode: Encoding-Based Protection against Bit-Flip Attacks on Neural Networks
by: Velčický, Patrik, et al.
Published: (2024) -
DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
by: Ou, Haoran, et al.
Published: (2026) -
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
by: Li, Yige, et al.
Published: (2025) -
A Critical Evaluation of Defenses against Prompt Injection Attacks
by: Jia, Yuqi, et al.
Published: (2025)