AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Chaohu, Gui, Tianyi, Liu, Yu, Xu, Linli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
von: Liu, Chenxi, et al.
Veröffentlicht: (2025)
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
von: Ji, Yuheng, et al.
Veröffentlicht: (2024)
von: Ji, Yuheng, et al.
Veröffentlicht: (2024)
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
von: Tao, Zhou, et al.
Veröffentlicht: (2025)
von: Tao, Zhou, et al.
Veröffentlicht: (2025)
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
von: Zhang, Hao, et al.
Veröffentlicht: (2024)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
von: Zhang, Ruohong, et al.
Veröffentlicht: (2024)
MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
von: Liu, Ziyu, et al.
Veröffentlicht: (2024)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
von: Tang, Jianting, et al.
Veröffentlicht: (2025)
von: Tang, Jianting, et al.
Veröffentlicht: (2025)
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2026)
Robust SAM: On the Adversarial Robustness of Vision Foundation Models
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
von: Long, Jiahuan, et al.
Veröffentlicht: (2025)
LPOI: Listwise Preference Optimization for Vision Language Models
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2025)
von: Zadeh, Fatemeh Pesaran, et al.
Veröffentlicht: (2025)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
von: Zhu, Kangyu, et al.
Veröffentlicht: (2024)
A Hybrid Defense Strategy for Boosting Adversarial Robustness in Vision-Language Models
von: Liang, Yuhan, et al.
Veröffentlicht: (2024)
von: Liang, Yuhan, et al.
Veröffentlicht: (2024)
NAP-Tuning: Neural Augmented Prompt Tuning for Adversarially Robust Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2025)
HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained Models
von: Liu, Han, et al.
Veröffentlicht: (2026)
von: Liu, Han, et al.
Veröffentlicht: (2026)
MIMIR: Masked Image Modeling for Mutual Information-based Adversarial Robustness
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2023)
von: Xu, Xiaoyun, et al.
Veröffentlicht: (2023)
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
von: Schlarmann, Christian, et al.
Veröffentlicht: (2024)
Navigating the Trade-off: A Synthesis of Defensive Strategies for Zero-Shot Adversarial Robustness in Vision-Language Models
von: Xu, Zane, et al.
Veröffentlicht: (2025)
von: Xu, Zane, et al.
Veröffentlicht: (2025)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
von: Wang, Yubo, et al.
Veröffentlicht: (2025)
Adversarial Prompt Distillation for Vision-Language Models
von: Luo, Lin, et al.
Veröffentlicht: (2024)
von: Luo, Lin, et al.
Veröffentlicht: (2024)
Hydra: An Agentic Reasoning Approach for Enhancing Adversarial Robustness and Mitigating Hallucinations in Vision-Language Models
von: Chung-En, et al.
Veröffentlicht: (2025)
von: Chung-En, et al.
Veröffentlicht: (2025)
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
von: Wu, Jiulong, et al.
Veröffentlicht: (2025)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhong, et al.
Veröffentlicht: (2026)
Probing the Robustness of Vision-Language Pretrained Models: A Multimodal Adversarial Attack Approach
von: Guan, Jiwei, et al.
Veröffentlicht: (2024)
von: Guan, Jiwei, et al.
Veröffentlicht: (2024)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
von: Li, Zhiwei, et al.
Veröffentlicht: (2025)
von: Li, Zhiwei, et al.
Veröffentlicht: (2025)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
von: Wu, Sihao, et al.
Veröffentlicht: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
von: Kim, Ju-Young, et al.
Veröffentlicht: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
ORCA: An Agentic Reasoning Framework for Hallucination and Adversarial Robustness in Vision-Language Models
von: Yu, Chung-En Johnny, et al.
Veröffentlicht: (2025)
von: Yu, Chung-En Johnny, et al.
Veröffentlicht: (2025)
PIP: Detecting Adversarial Examples in Large Vision-Language Models via Attention Patterns of Irrelevant Probe Questions
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
von: Zhang, Yudong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024) -
Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
von: Liu, Chenxi, et al.
Veröffentlicht: (2025) -
Enhancing Adversarial Robustness of Vision-Language Models through Low-Rank Adaptation
von: Ji, Yuheng, et al.
Veröffentlicht: (2024) -
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
von: Tao, Zhou, et al.
Veröffentlicht: (2025) -
B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions
von: Zhang, Hao, et al.
Veröffentlicht: (2024)