Enhancing Cross-Prompt Transferability in Vision-Language Models through Contextual Injection of Target Tokens
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Xikang, Tang, Xuehai, Zhu, Fuqing, Han, Jizhong, Hu, Songlin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
von: Yang, Xikang, et al.
Veröffentlicht: (2024)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023)
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
von: Wu, Junjie, et al.
Veröffentlicht: (2025)
von: Wu, Junjie, et al.
Veröffentlicht: (2025)
OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction
von: Wang, Xin, et al.
Veröffentlicht: (2026)
von: Wang, Xin, et al.
Veröffentlicht: (2026)
Self-Comparison for Dataset-Level Membership Inference in Large (Vision-)Language Models
von: Ren, Jie, et al.
Veröffentlicht: (2024)
von: Ren, Jie, et al.
Veröffentlicht: (2024)
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models for Medicine: A Comprehensive Survey
von: Ye, Jiarui, et al.
Veröffentlicht: (2025)
von: Ye, Jiarui, et al.
Veröffentlicht: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
von: Zhang, Yabin, et al.
Veröffentlicht: (2024)
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Federated Multi-Task Clustering
von: Dai, Suyan, et al.
Veröffentlicht: (2025)
von: Dai, Suyan, et al.
Veröffentlicht: (2025)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
von: Sun, Zeyi, et al.
Veröffentlicht: (2024)
DeepTextMark: A Deep Learning-Driven Text Watermarking Approach for Identifying Large Language Model Generated Text
von: Munyer, Travis, et al.
Veröffentlicht: (2023)
von: Munyer, Travis, et al.
Veröffentlicht: (2023)
FedMAC: Tackling Partial-Modality Missing in Federated Learning with Cross-Modal Aggregation and Contrastive Regularization
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2024)
von: Nguyen, Manh Duong, et al.
Veröffentlicht: (2024)
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
von: Qin, Yang, et al.
Veröffentlicht: (2025)
von: Qin, Yang, et al.
Veröffentlicht: (2025)
Efficient Distributed Training through Gradient Compression with Sparsification and Quantization Techniques
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
von: Singh, Shruti, et al.
Veröffentlicht: (2024)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
von: Pang, Xingzhou, et al.
Veröffentlicht: (2026)
QoS-QoE Translation with Large Language Model
von: Yu, Yingjie, et al.
Veröffentlicht: (2026)
von: Yu, Yingjie, et al.
Veröffentlicht: (2026)
Zero-shot image privacy classification with Vision-Language Models
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
von: Baia, Alina Elena, et al.
Veröffentlicht: (2025)
PTSBench: A Comprehensive Post-Training Sparsity Benchmark Towards Algorithms and Models
von: Wnag, Zining, et al.
Veröffentlicht: (2024)
von: Wnag, Zining, et al.
Veröffentlicht: (2024)
Multi-source Knowledge Enhanced Graph Attention Networks for Multimodal Fact Verification
von: Cao, Han, et al.
Veröffentlicht: (2024)
von: Cao, Han, et al.
Veröffentlicht: (2024)
Beyond Interpretability: Exploring the Comprehensibility of Adaptive Video Streaming through Large Language Models
von: Jia, Lianchen, et al.
Veröffentlicht: (2025)
von: Jia, Lianchen, et al.
Veröffentlicht: (2025)
Atom: Efficient On-Device Video-Language Pipelines Through Modular Reuse
von: Panchal, Kunjal, et al.
Veröffentlicht: (2025)
von: Panchal, Kunjal, et al.
Veröffentlicht: (2025)
Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
von: Farina, Matteo, et al.
Veröffentlicht: (2025)
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
von: Tong, Yijie, et al.
Veröffentlicht: (2026)
von: Tong, Yijie, et al.
Veröffentlicht: (2026)
Concept Drift Guided LayerNorm Tuning for Efficient Multimodal Metaphor Identification
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
von: Qian, Wenhao, et al.
Veröffentlicht: (2025)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
von: Zhu, Wentao
Veröffentlicht: (2024)
von: Zhu, Wentao
Veröffentlicht: (2024)
GUISE: Graph GaUssIan Shading watErmark
von: Yang, Renyi
Veröffentlicht: (2024)
von: Yang, Renyi
Veröffentlicht: (2024)
Modeling Musical Genre Trajectories through Pathlet Learning
von: Marey, Lilian, et al.
Veröffentlicht: (2025)
von: Marey, Lilian, et al.
Veröffentlicht: (2025)
Balanced Multimodal Learning: An Unidirectional Dynamic Interaction Perspective
von: Wang, Shijie, et al.
Veröffentlicht: (2025)
von: Wang, Shijie, et al.
Veröffentlicht: (2025)
RDPM: Solve Diffusion Probabilistic Models via Recurrent Token Prediction
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
von: Wu, Xiaoping, et al.
Veröffentlicht: (2024)
LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
von: Menon, Aditya Srinivas, et al.
Veröffentlicht: (2025)
Language Models as Black-Box Optimizers for Vision-Language Models
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning
von: Shen, Meng, et al.
Veröffentlicht: (2024)
von: Shen, Meng, et al.
Veröffentlicht: (2024)
Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
von: Pedretti, Carlo Teo, et al.
Veröffentlicht: (2025)
von: Pedretti, Carlo Teo, et al.
Veröffentlicht: (2025)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
von: Chen, Shengkai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
von: Yang, Xikang, et al.
Veröffentlicht: (2024) -
The Dark Side of Trust: Authority Citation-Driven Jailbreak Attacks on Large Language Models
von: Yang, Xikang, et al.
Veröffentlicht: (2024) -
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
von: Yang, Xikang, et al.
Veröffentlicht: (2025) -
Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2023) -
E2LVLM:Evidence-Enhanced Large Vision-Language Model for Multimodal Out-of-Context Misinformation Detection
von: Wu, Junjie, et al.
Veröffentlicht: (2025)