Saved in:
| Main Authors: | Wang, Xiaosen, Wang, Shaokang, Ge, Zhijin, Luo, Yuyang, Zhang, Shudong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.19911 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Security Risk of Misalignment between Text and Image in Multi-modal Model
by: Wang, Xiaosen, et al.
Published: (2025)
by: Wang, Xiaosen, et al.
Published: (2025)
Disrupting Semantic and Abstract Features for Better Adversarial Transferability
by: Luo, Yuyang, et al.
Published: (2025)
by: Luo, Yuyang, et al.
Published: (2025)
Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation
by: Wang, Xiaosen, et al.
Published: (2026)
by: Wang, Xiaosen, et al.
Published: (2026)
One Last Attention for Your Vision-Language Model
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)
by: Wang, Zhiqiang, et al.
Published: (2026)
Boosting the Local Invariance for Better Adversarial Transferability
by: Liu, Bohan, et al.
Published: (2025)
by: Liu, Bohan, et al.
Published: (2025)
ViT-EnsembleAttack: Augmenting Ensemble Models for Stronger Adversarial Transferability in Vision Transformers
by: Cao, Hanwen, et al.
Published: (2025)
by: Cao, Hanwen, et al.
Published: (2025)
IDEATOR: Jailbreaking and Benchmarking Large Vision-Language Models Using Themselves
by: Wang, Ruofan, et al.
Published: (2024)
by: Wang, Ruofan, et al.
Published: (2024)
Asymmetric Visual Semantic Embedding Framework for Efficient Vision-Language Alignment
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
by: Li, Chenxuan, et al.
Published: (2024)
by: Li, Chenxuan, et al.
Published: (2024)
SAFE-Pruner: Semantic Attention-Guided Future-Aware Token Pruning for Efficient Vision-Language-Action Manipulation
by: Ma, Shilin, et al.
Published: (2026)
by: Ma, Shilin, et al.
Published: (2026)
Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
PRIME: Protect Your Videos From Malicious Editing
by: Li, Guanlin, et al.
Published: (2024)
by: Li, Guanlin, et al.
Published: (2024)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
by: Fang, Zheng, et al.
Published: (2026)
by: Fang, Zheng, et al.
Published: (2026)
Disease-informed Adaptation of Vision-Language Models
by: Zhang, Jiajin, et al.
Published: (2024)
by: Zhang, Jiajin, et al.
Published: (2024)
AWT: Transferring Vision-Language Models via Augmentation, Weighting, and Transportation
by: Zhu, Yuhan, et al.
Published: (2024)
by: Zhu, Yuhan, et al.
Published: (2024)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
Bag of Tricks to Boost Adversarial Transferability
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
by: Zhang, Yi-Fan, et al.
Published: (2024)
by: Zhang, Yi-Fan, et al.
Published: (2024)
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
by: Xie, Chunlong, et al.
Published: (2025)
by: Xie, Chunlong, et al.
Published: (2025)
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
by: Zhang, Zhenhao, et al.
Published: (2026)
by: Zhang, Zhenhao, et al.
Published: (2026)
Your Vision-Language-Action Model Already Has Attention Heads For Path Deviation Detection
by: Jeong, Jaehwan, et al.
Published: (2026)
by: Jeong, Jaehwan, et al.
Published: (2026)
Less Could Be Better: Parameter-efficient Fine-tuning Advances Medical Vision Foundation Models
by: Lian, Chenyu, et al.
Published: (2024)
by: Lian, Chenyu, et al.
Published: (2024)
Malicious Path Manipulations via Exploitation of Representation Vulnerabilities of Vision-Language Navigation Systems
by: Islam, Chashi Mahiul, et al.
Published: (2024)
by: Islam, Chashi Mahiul, et al.
Published: (2024)
Improving Vision Transformers by Overlapping Heads in Multi-Head Self-Attention
by: Zhang, Tianxiao, et al.
Published: (2024)
by: Zhang, Tianxiao, et al.
Published: (2024)
Aligning Information Capacity Between Vision and Language via Dense-to-Sparse Feature Distillation for Image-Text Matching
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
RoboMamba: Efficient Vision-Language-Action Model for Robotic Reasoning and Manipulation
by: Liu, Jiaming, et al.
Published: (2024)
by: Liu, Jiaming, et al.
Published: (2024)
Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot
by: Hang, Sheng, et al.
Published: (2025)
by: Hang, Sheng, et al.
Published: (2025)
Diagnosing and Repairing Unsafe Channels in Vision-Language Models via Causal Discovery and Dual-Modal Safety Subspace Projection
by: Fu, Jinhu, et al.
Published: (2026)
by: Fu, Jinhu, et al.
Published: (2026)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
by: Wang, Hongyu, et al.
Published: (2025)
by: Wang, Hongyu, et al.
Published: (2025)
A Knowledge-guided Adversarial Defense for Resisting Malicious Visual Manipulation
by: Zhou, Dawei, et al.
Published: (2025)
by: Zhou, Dawei, et al.
Published: (2025)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
A-VL: Adaptive Attention for Large Vision-Language Models
by: Zhang, Junyang, et al.
Published: (2024)
by: Zhang, Junyang, et al.
Published: (2024)
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
by: Lu, Hannan, et al.
Published: (2024)
by: Lu, Hannan, et al.
Published: (2024)
Enhancing Vision-Language Models Generalization via Diversity-Driven Novel Feature Synthesis
by: Yan, Siyuan, et al.
Published: (2024)
by: Yan, Siyuan, et al.
Published: (2024)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
Large Vision-Language Models Get Lost in Attention
by: Xi, Gongli, et al.
Published: (2026)
by: Xi, Gongli, et al.
Published: (2026)
Dynamic Rank Adaptation for Vision-Language Models
by: Wang, Jiahui, et al.
Published: (2025)
by: Wang, Jiahui, et al.
Published: (2025)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Similar Items
-
Security Risk of Misalignment between Text and Image in Multi-modal Model
by: Wang, Xiaosen, et al.
Published: (2025) -
Disrupting Semantic and Abstract Features for Better Adversarial Transferability
by: Luo, Yuyang, et al.
Published: (2025) -
Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation
by: Wang, Xiaosen, et al.
Published: (2026) -
One Last Attention for Your Vision-Language Model
by: Chen, Liang, et al.
Published: (2025) -
Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models
by: Wang, Zhiqiang, et al.
Published: (2026)