Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Haoyu, Li, Dingcheng, Rutishauser, Lukas, Zheng, Zeyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Reinforcement Learning for Large Language Model Agent Safety
by: Wang, Zizhao, et al.
Published: (2025)
by: Wang, Zizhao, et al.
Published: (2025)
Reinforcement Learning with Backtracking Feedback
by: Sel, Bilgehan, et al.
Published: (2026)
by: Sel, Bilgehan, et al.
Published: (2026)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025)
by: Perin, Gabriel J., et al.
Published: (2025)
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
by: Shang, Yingjia, et al.
Published: (2025)
by: Shang, Yingjia, et al.
Published: (2025)
Cross-Modal Augmentation for Few-Shot Multimodal Fake News Detection
by: Jiang, Ye, et al.
Published: (2024)
by: Jiang, Ye, et al.
Published: (2024)
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
by: Jain, Yash, et al.
Published: (2024)
by: Jain, Yash, et al.
Published: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
by: Aldahoul, Nouar, et al.
Published: (2025)
by: Aldahoul, Nouar, et al.
Published: (2025)
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
by: Liu, Chenxi, et al.
Published: (2025)
by: Liu, Chenxi, et al.
Published: (2025)
Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training Approach
by: Jiang, Feihu, et al.
Published: (2024)
by: Jiang, Feihu, et al.
Published: (2024)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
by: Huang, James Y., et al.
Published: (2025)
by: Huang, James Y., et al.
Published: (2025)
Quantifying Modality Contributions via Disentangling Multimodal Representations
by: Amit, Padegal, et al.
Published: (2025)
by: Amit, Padegal, et al.
Published: (2025)
Safety Reasoning with Guidelines
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
Multimodal Contrastive Learning via Uni-Modal Coding and Cross-Modal Prediction for Multimodal Sentiment Analysis
by: Lin, Ronghao, et al.
Published: (2022)
by: Lin, Ronghao, et al.
Published: (2022)
A Depression Detection Method Based on Multi-Modal Feature Fusion Using Cross-Attention
by: Li, Shengjie, et al.
Published: (2024)
by: Li, Shengjie, et al.
Published: (2024)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
by: Mohanty, Shrestha, et al.
Published: (2024)
by: Mohanty, Shrestha, et al.
Published: (2024)
Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
by: Agarwal, Rajan, et al.
Published: (2025)
by: Agarwal, Rajan, et al.
Published: (2025)
Combating Adversarial Attacks with Multi-Agent Debate
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
Cross-Modal Consistency in Multimodal Large Language Models
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Enhancing Multimodal Sentiment Analysis for Missing Modality through Self-Distillation and Unified Modality Cross-Attention
by: Weng, Yuzhe, et al.
Published: (2024)
by: Weng, Yuzhe, et al.
Published: (2024)
MMSR: Symbolic Regression is a Multi-Modal Information Fusion Task
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
MobileCLIP2: Improving Multi-Modal Reinforced Training
by: Faghri, Fartash, et al.
Published: (2025)
by: Faghri, Fartash, et al.
Published: (2025)
Cross-Modal Navigation with Multi-Agent Reinforcement Learning
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
MMICT: Boosting Multi-Modal Fine-Tuning with In-Context Examples
by: Chen, Tao, et al.
Published: (2023)
by: Chen, Tao, et al.
Published: (2023)
Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Adversarial Training for Defense Against Label Poisoning Attacks
by: Bal, Melis Ilayda, et al.
Published: (2025)
by: Bal, Melis Ilayda, et al.
Published: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
by: Park, Jean, et al.
Published: (2024)
by: Park, Jean, et al.
Published: (2024)
Lightweight Safety Guardrails via Synthetic Data and RL-guided Adversarial Training
by: Ilin, Aleksei, et al.
Published: (2025)
by: Ilin, Aleksei, et al.
Published: (2025)
Towards Robust Alignment of Language Models: Distributionally Robustifying Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment
by: Liu, Chenxi, et al.
Published: (2024)
by: Liu, Chenxi, et al.
Published: (2024)
InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning
by: Zhu, Guanghao, et al.
Published: (2025)
by: Zhu, Guanghao, et al.
Published: (2025)
Fast Adversarial Training against Textual Adversarial Attacks
by: Yang, Yichen, et al.
Published: (2024)
by: Yang, Yichen, et al.
Published: (2024)
TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
by: Xing, Xiaobo, et al.
Published: (2025)
by: Xing, Xiaobo, et al.
Published: (2025)
WebInject: Prompt Injection Attack to Web Agents
by: Wang, Xilong, et al.
Published: (2025)
by: Wang, Xilong, et al.
Published: (2025)
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Multi-Modal Data Exploration via Language Agents
by: Nooralahzadeh, Farhad, et al.
Published: (2024)
by: Nooralahzadeh, Farhad, et al.
Published: (2024)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
by: Li, Kunxi, et al.
Published: (2025)
by: Li, Kunxi, et al.
Published: (2025)
Similar Items
-
Adversarial Reinforcement Learning for Large Language Model Agent Safety
by: Wang, Zizhao, et al.
Published: (2025) -
Reinforcement Learning with Backtracking Feedback
by: Sel, Bilgehan, et al.
Published: (2026) -
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
by: Perin, Gabriel J., et al.
Published: (2025) -
Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation
by: Shang, Yingjia, et al.
Published: (2025) -
Cross-Modal Augmentation for Few-Shot Multimodal Fake News Detection
by: Jiang, Ye, et al.
Published: (2024)