CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xu, Li, Hao, Lu, Zhichao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026)
by: Lin, Lixing, et al.
Published: (2026)
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026)
by: Li, Siyuan, et al.
Published: (2026)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024)
by: Li, Qinfeng, et al.
Published: (2024)
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
GuardReasoner: Towards Reasoning-based LLM Safeguards
by: Liu, Yue, et al.
Published: (2025)
by: Liu, Yue, et al.
Published: (2025)
PropGuard: Safeguarding LLM-MAS via Propagation-Aware Exploration and Remediation
by: Yan, Bingyu, et al.
Published: (2026)
by: Yan, Bingyu, et al.
Published: (2026)
Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection
by: Guo, Yihao, et al.
Published: (2025)
by: Guo, Yihao, et al.
Published: (2025)
DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization
by: Tang, Wenxin, et al.
Published: (2026)
by: Tang, Wenxin, et al.
Published: (2026)
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
by: Halawi, Danny, et al.
Published: (2024)
by: Halawi, Danny, et al.
Published: (2024)
LLM Safeguard is a Double-Edged Sword: Exploiting False Positives for Denial-of-Service Attacks
by: Zhang, Qingzhao, et al.
Published: (2024)
by: Zhang, Qingzhao, et al.
Published: (2024)
Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs
by: Yuan, Leitao, et al.
Published: (2026)
by: Yuan, Leitao, et al.
Published: (2026)
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders
by: Noever, David, et al.
Published: (2024)
by: Noever, David, et al.
Published: (2024)
CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024)
by: Li, Qinfeng, et al.
Published: (2024)
BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
by: Yan, Xiaobei, et al.
Published: (2025)
by: Yan, Xiaobei, et al.
Published: (2025)
MergeGuard: Efficient Thwarting of Trojan Attacks in Machine Learning Models
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2025)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
by: Ma, Jiachen, et al.
Published: (2024)
by: Ma, Jiachen, et al.
Published: (2024)
SAB:A Stealing and Robust Backdoor Attack based on Steganographic Algorithm against Federated Learning
by: Xu, Weida, et al.
Published: (2024)
by: Xu, Weida, et al.
Published: (2024)
VFEFL: Privacy-Preserving Federated Learning against Malicious Clients via Verifiable Functional Encryption
by: Cai, Nina, et al.
Published: (2025)
by: Cai, Nina, et al.
Published: (2025)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield Prompting
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
X-Guard: Multilingual Guard Agent for Content Moderation
by: Upadhayay, Bibek, et al.
Published: (2025)
by: Upadhayay, Bibek, et al.
Published: (2025)
FaceSwapGuard: Safeguarding Facial Privacy from DeepFake Threats through Identity Obfuscation
by: Wang, Li, et al.
Published: (2025)
by: Wang, Li, et al.
Published: (2025)
Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging
by: Li, Qinfeng, et al.
Published: (2025)
by: Li, Qinfeng, et al.
Published: (2025)
UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks
by: Yu, Tianlong, et al.
Published: (2026)
by: Yu, Tianlong, et al.
Published: (2026)
Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks
by: Xu, Yixiao, et al.
Published: (2025)
by: Xu, Yixiao, et al.
Published: (2025)
Safeguarding Text-to-Image Generative Models Against Unauthorized Knowledge Distillation
by: Gao, Yilan, et al.
Published: (2026)
by: Gao, Yilan, et al.
Published: (2026)
Safeguarding Large Language Models: A Survey
by: Dong, Yi, et al.
Published: (2024)
by: Dong, Yi, et al.
Published: (2024)
CUBA: Controlled Untargeted Backdoor Attack against Deep Neural Networks
by: Wu, Yinghao, et al.
Published: (2025)
by: Wu, Yinghao, et al.
Published: (2025)
PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems
by: Wang, Haozhen, et al.
Published: (2026)
by: Wang, Haozhen, et al.
Published: (2026)
Model Inversion Attack against Federated Unlearning
by: Zhou, Lei, et al.
Published: (2025)
by: Zhou, Lei, et al.
Published: (2025)
Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Generating Is Believing: Membership Inference Attacks against Retrieval-Augmented Generation
by: Li, Yuying, et al.
Published: (2024)
by: Li, Yuying, et al.
Published: (2024)
TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems
by: Wang, Kai, et al.
Published: (2026)
by: Wang, Kai, et al.
Published: (2026)
Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks
by: Fu, Haowei, et al.
Published: (2025)
by: Fu, Haowei, et al.
Published: (2025)
Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems
by: Wang, Haowei, et al.
Published: (2025)
by: Wang, Haowei, et al.
Published: (2025)
Explainer-guided Targeted Adversarial Attacks against Binary Code Similarity Detection Models
by: Chen, Mingjie, et al.
Published: (2025)
by: Chen, Mingjie, et al.
Published: (2025)
ShadowCode: Towards (Automatic) External Prompt Injection Attack against Code LLMs
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Persona Attack: Incremental Memory Injection Jailbreak Attack against Large Language Models
by: Park, Junyoung, et al.
Published: (2026)
by: Park, Junyoung, et al.
Published: (2026)
SoK: Robustness in Large Language Models against Jailbreak Attacks
by: Xu, Feiyue, et al.
Published: (2026)
by: Xu, Feiyue, et al.
Published: (2026)
Similar Items
-
Reflect-Guard: Enhancing LLM Safeguards against Adversarial Prompts via Logical Self-Reflection
by: Lin, Lixing, et al.
Published: (2026) -
CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks
by: Li, Siyuan, et al.
Published: (2026) -
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
by: Liu, Yue, et al.
Published: (2025) -
TransLinkGuard: Safeguarding Transformer Models Against Model Stealing in Edge Deployment
by: Li, Qinfeng, et al.
Published: (2024) -
SDD: Self-Degraded Defense against Malicious Fine-tuning
by: Chen, Zixuan, et al.
Published: (2025)