A Survey of Attacks on Large Vision-Language Models: Resources, Advances, and Future Trends
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Daizong, Yang, Mingyu, Qu, Xiaoye, Zhou, Pan, Cheng, Yu, Hu, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
von: Liu, Daizong, et al.
Veröffentlicht: (2024)
von: Liu, Daizong, et al.
Veröffentlicht: (2024)
Rethinking Video-Language Model from the Language Input Perspective
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
von: Song, Mingyang, et al.
Veröffentlicht: (2025)
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
SURf: Teaching Large Vision-Language Models to Selectively Utilize Retrieved Information
von: Sun, Jiashuo, et al.
Veröffentlicht: (2024)
von: Sun, Jiashuo, et al.
Veröffentlicht: (2024)
VideoSSR: Video Self-Supervised Reinforcement Learning
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding
von: Huang, Wencan, et al.
Veröffentlicht: (2025)
von: Huang, Wencan, et al.
Veröffentlicht: (2025)
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
Spotlight on Token Perception for Multimodal Reinforcement Learning
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
von: Huang, Siyuan, et al.
Veröffentlicht: (2025)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
von: Luo, Zhi, et al.
Veröffentlicht: (2025)
von: Luo, Zhi, et al.
Veröffentlicht: (2025)
Improving the Transferability of 3D Point Cloud Attack via Spectral-aware Admix and Optimization Designs
von: Hu, Shiyu, et al.
Veröffentlicht: (2024)
von: Hu, Shiyu, et al.
Veröffentlicht: (2024)
Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval Using Language
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Alleviating Hallucination in Large Vision-Language Models with Active Retrieval Augmentation
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
Hard-Label Black-Box Attacks on 3D Point Clouds
von: Liu, Daizong, et al.
Veröffentlicht: (2024)
von: Liu, Daizong, et al.
Veröffentlicht: (2024)
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Vision Language Models in Autonomous Driving: A Survey and Outlook
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2023)
From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy
von: Du, Bi'an, et al.
Veröffentlicht: (2026)
von: Du, Bi'an, et al.
Veröffentlicht: (2026)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
von: Ye, Mang, et al.
Veröffentlicht: (2025)
von: Ye, Mang, et al.
Veröffentlicht: (2025)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
von: Shen, Chuming, et al.
Veröffentlicht: (2025)
von: Shen, Chuming, et al.
Veröffentlicht: (2025)
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
Physical Prompt Injection Attacks on Large Vision-Language Models
von: Ling, Chen, et al.
Veröffentlicht: (2026)
von: Ling, Chen, et al.
Veröffentlicht: (2026)
Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Stealthy Backdoor Attack in Self-Supervised Learning Vision Encoders for Large Vision Language Models
von: Liu, Zhaoyi, et al.
Veröffentlicht: (2025)
von: Liu, Zhaoyi, et al.
Veröffentlicht: (2025)
Multi-Turn Adaptive Prompting Attack on Large Vision-Language Models
von: Choi, In Chong, et al.
Veröffentlicht: (2026)
von: Choi, In Chong, et al.
Veröffentlicht: (2026)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
von: Zhang, Zhifang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhifang, et al.
Veröffentlicht: (2025)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
von: Fang, Xiang, et al.
Veröffentlicht: (2022)
Towards Vision-Language Geo-Foundation Model: A Survey
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
von: Zhou, Yue, et al.
Veröffentlicht: (2024)
Black-Box Adversarial Attack on Vision Language Models for Autonomous Driving
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
Mitigating Information Loss under High Pruning Rates for Efficient Large Vision Language Models
von: Fu, Mingyu, et al.
Veröffentlicht: (2025)
von: Fu, Mingyu, et al.
Veröffentlicht: (2025)
Visual Adversarial Attack on Vision-Language Models for Autonomous Driving
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyuan, et al.
Veröffentlicht: (2024)
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
von: Chen, Guanjie, et al.
Veröffentlicht: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
von: Tian, Jie, et al.
Veröffentlicht: (2025)
von: Tian, Jie, et al.
Veröffentlicht: (2025)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Survey on Text-guided 3D Visual Grounding: Elements, Recent Advances, and Future Directions
von: Liu, Daizong, et al.
Veröffentlicht: (2024) -
Rethinking Video-Language Model from the Language Input Perspective
von: Fang, Xiang, et al.
Veröffentlicht: (2026) -
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024) -
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
von: Song, Mingyang, et al.
Veröffentlicht: (2025) -
Mitigating Multilingual Hallucination in Large Vision-Language Models
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)