Dissecting Adversarial Robustness of Multimodal LM Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Chen Henry, Shah, Rishi, Koh, Jing Yu, Salakhutdinov, Ruslan, Fried, Daniel, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025)
by: Shah, Rishi Rajesh, et al.
Published: (2025)
DeepSight: An All-in-One LM Safety Toolkit
by: Zhang, Bo, et al.
Published: (2026)
by: Zhang, Bo, et al.
Published: (2026)
Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
by: Qu, Jingen, et al.
Published: (2025)
by: Qu, Jingen, et al.
Published: (2025)
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
by: Xiong, Yuan, et al.
Published: (2025)
by: Xiong, Yuan, et al.
Published: (2025)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
by: Huang, Youcheng, et al.
Published: (2024)
by: Huang, Youcheng, et al.
Published: (2024)
Doubly-Universal Adversarial Perturbations: Deceiving Vision-Language Models Across Both Images and Text with a Single Perturbation
by: Kim, Hee-Seon, et al.
Published: (2024)
by: Kim, Hee-Seon, et al.
Published: (2024)
Iteratively Prompting Multimodal LLMs to Reproduce Natural and AI-Generated Images
by: Naseh, Ali, et al.
Published: (2024)
by: Naseh, Ali, et al.
Published: (2024)
VLSBench: Unveiling Visual Leakage in Multimodal Safety
by: Hu, Xuhao, et al.
Published: (2024)
by: Hu, Xuhao, et al.
Published: (2024)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
by: Kim, Taeyoun, et al.
Published: (2024)
by: Kim, Taeyoun, et al.
Published: (2024)
Rethinking Robust Adversarial Concept Erasure in Diffusion Models
by: Yin, Qinghong, et al.
Published: (2025)
by: Yin, Qinghong, et al.
Published: (2025)
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
by: Yang, Wenkui, et al.
Published: (2026)
by: Yang, Wenkui, et al.
Published: (2026)
Navigation as Attackers Wish? Towards Building Robust Embodied Agents under Federated Learning
by: Zhang, Yunchao, et al.
Published: (2022)
by: Zhang, Yunchao, et al.
Published: (2022)
Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
by: Miao, Ziqi, et al.
Published: (2025)
by: Miao, Ziqi, et al.
Published: (2025)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
by: Ding, Yi, et al.
Published: (2025)
by: Ding, Yi, et al.
Published: (2025)
Multi-Agent Computer Use
by: Koh, Jing Yu, et al.
Published: (2026)
by: Koh, Jing Yu, et al.
Published: (2026)
Improving Adversarial Robustness via Feature Pattern Consistency Constraint
by: Hu, Jiacong, et al.
Published: (2024)
by: Hu, Jiacong, et al.
Published: (2024)
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
by: Gu, Xiangming, et al.
Published: (2024)
by: Gu, Xiangming, et al.
Published: (2024)
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition
by: Fioresi, Joseph, et al.
Published: (2025)
by: Fioresi, Joseph, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Is RobustBench/AutoAttack a suitable Benchmark for Adversarial Robustness?
by: Lorenz, Peter, et al.
Published: (2021)
by: Lorenz, Peter, et al.
Published: (2021)
On the Robustness of Kolmogorov-Arnold Networks: An Adversarial Perspective
by: Alter, Tal, et al.
Published: (2024)
by: Alter, Tal, et al.
Published: (2024)
Distilling Adversarial Robustness Using Heterogeneous Teachers
by: Deng, Jieren, et al.
Published: (2024)
by: Deng, Jieren, et al.
Published: (2024)
$\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models
by: Wu, Huanqi, et al.
Published: (2025)
by: Wu, Huanqi, et al.
Published: (2025)
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
by: Tian, Yuan, et al.
Published: (2026)
by: Tian, Yuan, et al.
Published: (2026)
ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial Optimization
by: Huang, Huayang, et al.
Published: (2024)
by: Huang, Huayang, et al.
Published: (2024)
PatchCURE: Improving Certifiable Robustness, Model Utility, and Computation Efficiency of Adversarial Patch Defenses
by: Xiang, Chong, et al.
Published: (2023)
by: Xiang, Chong, et al.
Published: (2023)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
by: Zhou, Yiwei, et al.
Published: (2024)
by: Zhou, Yiwei, et al.
Published: (2024)
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
by: Lee, DongGeon, et al.
Published: (2025)
by: Lee, DongGeon, et al.
Published: (2025)
Deciphering the Definition of Adversarial Robustness for post-hoc OOD Detectors
by: Lorenz, Peter, et al.
Published: (2024)
by: Lorenz, Peter, et al.
Published: (2024)
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators
by: Tian, Yu, et al.
Published: (2024)
by: Tian, Yu, et al.
Published: (2024)
RoMA: Robust Malware Attribution via Byte-level Adversarial Training with Global Perturbations and Adversarial Consistency Regularization
by: Sun, Yuxia, et al.
Published: (2025)
by: Sun, Yuxia, et al.
Published: (2025)
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models
by: Zhang, Yimeng, et al.
Published: (2024)
by: Zhang, Yimeng, et al.
Published: (2024)
Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents
by: Jha, Rishi, et al.
Published: (2026)
by: Jha, Rishi, et al.
Published: (2026)
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
by: Ma, Xingjun, et al.
Published: (2025)
by: Ma, Xingjun, et al.
Published: (2025)
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
by: Yin, Ziyi, et al.
Published: (2023)
by: Yin, Ziyi, et al.
Published: (2023)
Robust Image Classification: Defensive Strategies against FGSM and PGD Adversarial Attacks
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
A Random Ensemble of Encrypted models for Enhancing Robustness against Adversarial Examples
by: Iijima, Ryota, et al.
Published: (2024)
by: Iijima, Ryota, et al.
Published: (2024)
Jailbreaking Attack against Multimodal Large Language Model
by: Niu, Zhenxing, et al.
Published: (2024)
by: Niu, Zhenxing, et al.
Published: (2024)
Boosting Generative Adversarial Transferability with Self-supervised Vision Transformer Features
by: Wu, Shangbo, et al.
Published: (2025)
by: Wu, Shangbo, et al.
Published: (2025)
Robust Watermarks Leak: Channel-Aware Feature Extraction Enables Adversarial Watermark Manipulation
by: Ba, Zhongjie, et al.
Published: (2025)
by: Ba, Zhongjie, et al.
Published: (2025)
Similar Items
-
Jailbreaking in the Haystack
by: Shah, Rishi Rajesh, et al.
Published: (2025) -
DeepSight: An All-in-One LM Safety Toolkit
by: Zhang, Bo, et al.
Published: (2026) -
Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
by: Qu, Jingen, et al.
Published: (2025) -
Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
by: Xiong, Yuan, et al.
Published: (2025) -
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
by: Huang, Youcheng, et al.
Published: (2024)