On the Geometric Limits of Transformer Defenses against Obfuscation Attacks: Latent Embedding Collapse & Performance Robustness Gap
Fuente:
arXiv
Guardado en:
| Autores principales: | Mashaido, Becky, Das, Tapadhir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024)
por: Xiong, Chen, et al.
Publicado: (2024)
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
por: Li, Shaofei, et al.
Publicado: (2024)
por: Li, Shaofei, et al.
Publicado: (2024)
PhantomFetch: Obfuscating Loads against Prefetcher Side-Channel Attacks
por: Zhang, Xingzhi, et al.
Publicado: (2025)
por: Zhang, Xingzhi, et al.
Publicado: (2025)
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
por: Liu, Deng, et al.
Publicado: (2026)
por: Liu, Deng, et al.
Publicado: (2026)
Removal Attack and Defense on AI-generated Content Latent-based Watermarking
por: Lee, De Zhang, et al.
Publicado: (2025)
por: Lee, De Zhang, et al.
Publicado: (2025)
Defense against Poisoning Attacks under Shuffle-DP
por: Wang, Siyi, et al.
Publicado: (2026)
por: Wang, Siyi, et al.
Publicado: (2026)
Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks
por: Tsiolakis, Theodoros, et al.
Publicado: (2025)
por: Tsiolakis, Theodoros, et al.
Publicado: (2025)
Mirage: Defense against CrossPath Attacks in Software Defined Networks
por: Murtuza, Shariq, et al.
Publicado: (2024)
por: Murtuza, Shariq, et al.
Publicado: (2024)
A Data-Driven Defense against Edge-case Model Poisoning Attacks on Federated Learning
por: Purohit, Kiran, et al.
Publicado: (2023)
por: Purohit, Kiran, et al.
Publicado: (2023)
System Prompt Extraction Attacks and Defenses in Large Language Models
por: Das, Badhan Chandra, et al.
Publicado: (2025)
por: Das, Badhan Chandra, et al.
Publicado: (2025)
From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks
por: Kulkarni, Aditya, et al.
Publicado: (2024)
por: Kulkarni, Aditya, et al.
Publicado: (2024)
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
por: Zhang, Kaiyuan, et al.
Publicado: (2025)
Obfuscating Code Vulnerabilities against Static Analysis in JavaScript Code
por: Pagano, Francesco, et al.
Publicado: (2026)
por: Pagano, Francesco, et al.
Publicado: (2026)
System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective
por: Wu, Fangzhou, et al.
Publicado: (2024)
por: Wu, Fangzhou, et al.
Publicado: (2024)
DAPPER: A Performance-Attack-Resilient Tracker for RowHammer Defense
por: Woo, Jeonghyun, et al.
Publicado: (2025)
por: Woo, Jeonghyun, et al.
Publicado: (2025)
A Critical Evaluation of Defenses against Prompt Injection Attacks
por: Jia, Yuqi, et al.
Publicado: (2025)
por: Jia, Yuqi, et al.
Publicado: (2025)
Robust Image Classification: Defensive Strategies against FGSM and PGD Adversarial Attacks
por: Waghela, Hetvi, et al.
Publicado: (2024)
por: Waghela, Hetvi, et al.
Publicado: (2024)
Empirical Assessment of the Code Comprehension Effort Needed to Attack Programs Protected with Obfuscation
por: Regano, Leonardo, et al.
Publicado: (2025)
por: Regano, Leonardo, et al.
Publicado: (2025)
Prototype-Guided Robust Learning against Backdoor Attacks
por: Guo, Wei, et al.
Publicado: (2025)
por: Guo, Wei, et al.
Publicado: (2025)
Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models
por: Truong, Vu Tuan, et al.
Publicado: (2026)
por: Truong, Vu Tuan, et al.
Publicado: (2026)
On Feasibility of Intent Obfuscating Attacks
por: Li, Zhaobin, et al.
Publicado: (2024)
por: Li, Zhaobin, et al.
Publicado: (2024)
HOACS: Homomorphic Obfuscation Assisted Concealing of Secrets to Thwart Trojan Attacks in COTS Processor
por: Hossain, Tanvir, et al.
Publicado: (2024)
por: Hossain, Tanvir, et al.
Publicado: (2024)
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
por: Zeng, Yifan, et al.
Publicado: (2024)
por: Zeng, Yifan, et al.
Publicado: (2024)
Towards Robust IoT Defense: Comparative Statistics of Attack Detection in Resource-Constrained Scenarios
por: Alwaisi, Zainab, et al.
Publicado: (2024)
por: Alwaisi, Zainab, et al.
Publicado: (2024)
Shield Bash: Abusing Defensive Coherence State Retrieval to Break Timing Obfuscation
por: Ramkrishnan, Kartik, et al.
Publicado: (2025)
por: Ramkrishnan, Kartik, et al.
Publicado: (2025)
BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
por: Gibert, Daniel, et al.
Publicado: (2024)
por: Gibert, Daniel, et al.
Publicado: (2024)
CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation
por: Yang, Xin, et al.
Publicado: (2025)
por: Yang, Xin, et al.
Publicado: (2025)
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
por: Chen, Yulin, et al.
Publicado: (2024)
por: Chen, Yulin, et al.
Publicado: (2024)
ModelObfuscator: Obfuscating Model Information to Protect Deployed ML-based Systems
por: Zhou, Mingyi, et al.
Publicado: (2023)
por: Zhou, Mingyi, et al.
Publicado: (2023)
ModelShield: Adaptive and Robust Watermark against Model Extraction Attack
por: Pang, Kaiyi, et al.
Publicado: (2024)
por: Pang, Kaiyi, et al.
Publicado: (2024)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
por: Fu, Haowei, et al.
Publicado: (2025)
por: Fu, Haowei, et al.
Publicado: (2025)
Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning
por: Luo, Xinjian, et al.
Publicado: (2020)
por: Luo, Xinjian, et al.
Publicado: (2020)
Defense against Joint Poison and Evasion Attacks: A Case Study of DERMS
por: Abdeen, Zain ul, et al.
Publicado: (2024)
por: Abdeen, Zain ul, et al.
Publicado: (2024)
Inferring Sensitive Attributes from Knowledge Graph Embeddings: Attack and Defense Strategies
por: Hayder, Yasmine
Publicado: (2026)
por: Hayder, Yasmine
Publicado: (2026)
Federated Learning: Attacks, Defenses, Opportunities, and Challenges
por: Shirvani, Ghazaleh, et al.
Publicado: (2024)
por: Shirvani, Ghazaleh, et al.
Publicado: (2024)
System Password Security: Attack and Defense Mechanisms
por: Shi, Chaofang, et al.
Publicado: (2025)
por: Shi, Chaofang, et al.
Publicado: (2025)
Provably Robust Explainable Graph Neural Networks against Graph Perturbation Attacks
por: Li, Jiate, et al.
Publicado: (2025)
por: Li, Jiate, et al.
Publicado: (2025)
Disassembling Obfuscated Executables with LLM
por: Rong, Huanyao, et al.
Publicado: (2024)
por: Rong, Huanyao, et al.
Publicado: (2024)
SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition
por: Luo, Chenxiang, et al.
Publicado: (2025)
por: Luo, Chenxiang, et al.
Publicado: (2025)
Ejemplares similares
-
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak Attacks
por: Xiong, Chen, et al.
Publicado: (2024) -
Query Provenance Analysis: Efficient and Robust Defense against Query-based Black-box Attacks
por: Li, Shaofei, et al.
Publicado: (2024) -
PhantomFetch: Obfuscating Loads against Prefetcher Side-Channel Attacks
por: Zhang, Xingzhi, et al.
Publicado: (2025) -
Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models
por: Liu, Deng, et al.
Publicado: (2026) -
Removal Attack and Defense on AI-generated Content Latent-based Watermarking
por: Lee, De Zhang, et al.
Publicado: (2025)