One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhao, Jiale, Jiang, Xinyang, Gao, Junyao, Xue, Yuhao, Zhao, Cairong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
di: Xue, Yuhao, et al.
Pubblicazione: (2025)
di: Xue, Yuhao, et al.
Pubblicazione: (2025)
DiffPhysBA: Diffusion-based Physical Backdoor Attack against Person Re-Identification in Real-World
di: Sun, Wenli, et al.
Pubblicazione: (2024)
di: Sun, Wenli, et al.
Pubblicazione: (2024)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
di: Wang, Yubin, et al.
Pubblicazione: (2024)
di: Wang, Yubin, et al.
Pubblicazione: (2024)
Similarity Distribution based Membership Inference Attack on Person Re-identification
di: Gao, Junyao, et al.
Pubblicazione: (2022)
di: Gao, Junyao, et al.
Pubblicazione: (2022)
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
di: Wang, Yubin, et al.
Pubblicazione: (2025)
di: Wang, Yubin, et al.
Pubblicazione: (2025)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
di: Ji, Chenhao, et al.
Pubblicazione: (2025)
di: Ji, Chenhao, et al.
Pubblicazione: (2025)
AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models
di: Li, Jiayu, et al.
Pubblicazione: (2025)
di: Li, Jiayu, et al.
Pubblicazione: (2025)
A Unified Understanding of Adversarial Vulnerability Regarding Unimodal Models and Vision-Language Pre-training Models
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
DROP: Decouple Re-Identification and Human Parsing with Task-specific Features for Occluded Person Re-identification
di: Dou, Shuguang, et al.
Pubblicazione: (2024)
di: Dou, Shuguang, et al.
Pubblicazione: (2024)
Sample-agnostic Adversarial Perturbation for Vision-Language Pre-training Models
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
di: Zheng, Haonan, et al.
Pubblicazione: (2024)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
di: Wang, Yubin, et al.
Pubblicazione: (2024)
di: Wang, Yubin, et al.
Pubblicazione: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
di: Song, Jiale, et al.
Pubblicazione: (2026)
di: Song, Jiale, et al.
Pubblicazione: (2026)
FaceShot: Bring Any Character into Life
di: Gao, Junyao, et al.
Pubblicazione: (2025)
di: Gao, Junyao, et al.
Pubblicazione: (2025)
StyleShot: A Snapshot on Any Style
di: Gao, Junyao, et al.
Pubblicazione: (2024)
di: Gao, Junyao, et al.
Pubblicazione: (2024)
Online Video Quality Enhancement with Spatial-Temporal Look-up Tables
di: Qu, Zefan, et al.
Pubblicazione: (2023)
di: Qu, Zefan, et al.
Pubblicazione: (2023)
PapMOT: Exploring Adversarial Patch Attack against Multiple Object Tracking
di: Long, Jiahuan, et al.
Pubblicazione: (2025)
di: Long, Jiahuan, et al.
Pubblicazione: (2025)
CharacterShot: Controllable and Consistent 4D Character Animation
di: Gao, Junyao, et al.
Pubblicazione: (2025)
di: Gao, Junyao, et al.
Pubblicazione: (2025)
Boosting Adversarial Transferability via Commonality-Oriented Gradient Optimization
di: Gao, Yanting, et al.
Pubblicazione: (2025)
di: Gao, Yanting, et al.
Pubblicazione: (2025)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
di: Zhang, Tianyuan, et al.
Pubblicazione: (2024)
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
di: Waseda, Futa, et al.
Pubblicazione: (2025)
di: Waseda, Futa, et al.
Pubblicazione: (2025)
BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
di: Li, Juncheng, et al.
Pubblicazione: (2025)
di: Li, Juncheng, et al.
Pubblicazione: (2025)
Exploring the Limits of Vision-Language-Action Manipulations in Cross-task Generalization
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
di: Zhou, Jiaming, et al.
Pubblicazione: (2025)
A Unified Perspective on Adversarial Membership Manipulation in Vision Models
di: Gao, Ruize, et al.
Pubblicazione: (2026)
di: Gao, Ruize, et al.
Pubblicazione: (2026)
Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
di: Xie, Peng, et al.
Pubblicazione: (2024)
di: Xie, Peng, et al.
Pubblicazione: (2024)
OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models
di: Zhang, Yiwei, et al.
Pubblicazione: (2026)
di: Zhang, Yiwei, et al.
Pubblicazione: (2026)
Adversarial Prompt Distillation for Vision-Language Models
di: Luo, Lin, et al.
Pubblicazione: (2024)
di: Luo, Lin, et al.
Pubblicazione: (2024)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
di: Luo, Haochen, et al.
Pubblicazione: (2024)
di: Luo, Haochen, et al.
Pubblicazione: (2024)
Benchmarking Vision-Language Models under Contradictory Virtual Content Attacks in Augmented Reality
di: Xiu, Yanming, et al.
Pubblicazione: (2026)
di: Xiu, Yanming, et al.
Pubblicazione: (2026)
PG-Attack: A Precision-Guided Adversarial Attack Framework Against Vision Foundation Models for Autonomous Driving
di: Fu, Jiyuan, et al.
Pubblicazione: (2024)
di: Fu, Jiyuan, et al.
Pubblicazione: (2024)
AGC: Adaptive Geodesic Correction for Adversarial Robustness on Vision-Language Models
di: Li, Zhiwei, et al.
Pubblicazione: (2026)
di: Li, Zhiwei, et al.
Pubblicazione: (2026)
MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
di: Gao, Junyao, et al.
Pubblicazione: (2026)
di: Gao, Junyao, et al.
Pubblicazione: (2026)
MAA: Meticulous Adversarial Attack against Vision-Language Pre-trained Models
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2025)
di: Zhang, Peng-Fei, et al.
Pubblicazione: (2025)
Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack
di: Jia, Xiaojun, et al.
Pubblicazione: (2024)
di: Jia, Xiaojun, et al.
Pubblicazione: (2024)
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
di: Shi, Zhonghao, et al.
Pubblicazione: (2025)
di: Shi, Zhonghao, et al.
Pubblicazione: (2025)
When and Where to Attack? Stage-wise Attention-Guided Adversarial Attack on Large Vision Language Models
di: Kwak, Jaehyun, et al.
Pubblicazione: (2026)
di: Kwak, Jaehyun, et al.
Pubblicazione: (2026)
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
Text-promptable Object Counting via Quantity Awareness Enhancement
di: Shi, Miaojing, et al.
Pubblicazione: (2025)
di: Shi, Miaojing, et al.
Pubblicazione: (2025)
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models
di: Hu, Chengyin, et al.
Pubblicazione: (2026)
di: Hu, Chengyin, et al.
Pubblicazione: (2026)
Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation
di: Gao, Zhenghao, et al.
Pubblicazione: (2024)
di: Gao, Zhenghao, et al.
Pubblicazione: (2024)
Transferable Adversarial Attacks on Black-Box Vision-Language Models
di: Hu, Kai, et al.
Pubblicazione: (2025)
di: Hu, Kai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
di: Xue, Yuhao, et al.
Pubblicazione: (2025) -
DiffPhysBA: Diffusion-based Physical Backdoor Attack against Person Re-Identification in Real-World
di: Sun, Wenli, et al.
Pubblicazione: (2024) -
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling
di: Wang, Yubin, et al.
Pubblicazione: (2024) -
Similarity Distribution based Membership Inference Attack on Person Re-identification
di: Gao, Junyao, et al.
Pubblicazione: (2022) -
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
di: Wang, Yubin, et al.
Pubblicazione: (2025)