Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Kuofeng, Gu, Jindong, Bai, Yang, Xia, Shu-Tao, Torr, Philip, Liu, Wei, Li, Zhifeng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023)
by: Bai, Jiawang, et al.
Published: (2023)
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
An Image Is Worth 1000 Lies: Adversarial Transferability across Prompts on Vision-Language Models
by: Luo, Haochen, et al.
Published: (2024)
by: Luo, Haochen, et al.
Published: (2024)
Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models
by: Gao, Kuofeng, et al.
Published: (2025)
by: Gao, Kuofeng, et al.
Published: (2025)
Video Watermarking: Safeguarding Your Video from (Unauthorized) Annotations by Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024)
by: Li, Jinmin, et al.
Published: (2024)
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
by: Chen, Haokun, et al.
Published: (2025)
by: Chen, Haokun, et al.
Published: (2025)
Influencer Backdoor Attack on Semantic Segmentation
by: Lan, Haoheng, et al.
Published: (2023)
by: Lan, Haoheng, et al.
Published: (2023)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
by: Wu, Yixuan, et al.
Published: (2025)
by: Wu, Yixuan, et al.
Published: (2025)
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
by: Liu, Haitong, et al.
Published: (2025)
by: Liu, Haitong, et al.
Published: (2025)
Improving Adversarial Transferability via Model Alignment
by: Ma, Avery, et al.
Published: (2023)
by: Ma, Avery, et al.
Published: (2023)
Self-Discovering Interpretable Diffusion Latent Directions for Responsible Text-to-Image Generation
by: Li, Hang, et al.
Published: (2023)
by: Li, Hang, et al.
Published: (2023)
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?
by: Hu, Anjun, et al.
Published: (2024)
by: Hu, Anjun, et al.
Published: (2024)
Which Model Generated This Image? A Model-Agnostic Approach for Origin Attribution
by: Liu, Fengyuan, et al.
Published: (2024)
by: Liu, Fengyuan, et al.
Published: (2024)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Seeing Through the Chain: Mitigate Hallucination in Multimodal Reasoning Models via CoT Compression and Contrastive Preference Optimization
by: Fang, Hao, et al.
Published: (2026)
by: Fang, Hao, et al.
Published: (2026)
Not All Prompts Are Secure: A Switchable Backdoor Attack Against Pre-trained Vision Transformers
by: Yang, Sheng, et al.
Published: (2024)
by: Yang, Sheng, et al.
Published: (2024)
When LLMs step into the 3D World: A Survey and Meta-Analysis of 3D Tasks via Multi-modal Large Language Models
by: Ma, Xianzheng, et al.
Published: (2024)
by: Ma, Xianzheng, et al.
Published: (2024)
Finetune Like You Pretrain: Boosting Zero-shot Adversarial Robustness in Vision-language Models
by: Xing, Songlong, et al.
Published: (2026)
by: Xing, Songlong, et al.
Published: (2026)
LLMRA: Multi-modal Large Language Model based Restoration Assistant
by: Jin, Xiaoyu, et al.
Published: (2024)
by: Jin, Xiaoyu, et al.
Published: (2024)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
by: Chen, Shuo, et al.
Published: (2023)
by: Chen, Shuo, et al.
Published: (2023)
Empowering Segmentation Ability to Multi-modal Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
Manipulation Facing Threats: Evaluating Physical Vulnerabilities in End-to-End Vision Language Action Models
by: Cheng, Hao, et al.
Published: (2024)
by: Cheng, Hao, et al.
Published: (2024)
True Multimodal In-Context Learning Needs Attention to the Visual Context
by: Chen, Shuo, et al.
Published: (2025)
by: Chen, Shuo, et al.
Published: (2025)
An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
by: Luo, Zhi, et al.
Published: (2025)
by: Luo, Zhi, et al.
Published: (2025)
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
by: Wen, Zichen, et al.
Published: (2025)
by: Wen, Zichen, et al.
Published: (2025)
Latent Guard: a Safety Framework for Text-to-image Generation
by: Liu, Runtao, et al.
Published: (2024)
by: Liu, Runtao, et al.
Published: (2024)
Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
by: Fang, Hao, et al.
Published: (2025)
by: Fang, Hao, et al.
Published: (2025)
Neural Antidote: Class-Wise Prompt Tuning for Purifying Backdoors in CLIP
by: Kong, Jiawei, et al.
Published: (2025)
by: Kong, Jiawei, et al.
Published: (2025)
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
by: Peng, Zelin, et al.
Published: (2025)
by: Peng, Zelin, et al.
Published: (2025)
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
by: Gu, Chenhe, et al.
Published: (2025)
by: Gu, Chenhe, et al.
Published: (2025)
DIPR: Efficient Point Cloud Registration via Dynamic Iteration
by: Ai, Yang, et al.
Published: (2023)
by: Ai, Yang, et al.
Published: (2023)
Localizing Events in Videos with Multimodal Queries
by: Zhang, Gengyuan, et al.
Published: (2024)
by: Zhang, Gengyuan, et al.
Published: (2024)
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
by: Liu, Xin, et al.
Published: (2023)
by: Liu, Xin, et al.
Published: (2023)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
by: Wu, Junfeng, et al.
Published: (2024)
by: Wu, Junfeng, et al.
Published: (2024)
Fine-grained Multiple Supervisory Network for Multi-modal Manipulation Detecting and Grounding
by: Yu, Xinquan, et al.
Published: (2025)
by: Yu, Xinquan, et al.
Published: (2025)
Similar Items
-
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024) -
Adversarial Robustness for Visual Grounding of Multimodal Large Language Models
by: Gao, Kuofeng, et al.
Published: (2024) -
BadCLIP: Trigger-Aware Prompt Learning for Backdoor Attacks on CLIP
by: Bai, Jiawang, et al.
Published: (2023) -
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
by: Li, Jinmin, et al.
Published: (2024) -
Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
by: Gao, Kuofeng, et al.
Published: (2024)