Explaining multimodal LLMs via intra-modal token interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Jiawei, Chen, Ruoyu, Jiao, Xianghao, Liang, Siyuan, Liu, Shiming, Zhang, Qunli, Hu, Zheng, Cao, Xiaochun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
by: Chen, Yannan, et al.
Published: (2025)
by: Chen, Yannan, et al.
Published: (2025)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Less is More: Fewer Interpretable Region via Submodular Subset Selection
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Adversarial Backdoor Defense in CLIP
by: Kuang, Junhao, et al.
Published: (2024)
by: Kuang, Junhao, et al.
Published: (2024)
R-PGA: Robust Physical Adversarial Camouflage Generation via Relightable 3D Gaussian Splatting
by: Lou, Tianrui, et al.
Published: (2026)
by: Lou, Tianrui, et al.
Published: (2026)
Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery Detection
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
by: Liang, Jiawei, et al.
Published: (2022)
by: Liang, Jiawei, et al.
Published: (2022)
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making
by: Chen, Ruoyu, et al.
Published: (2026)
by: Chen, Ruoyu, et al.
Published: (2026)
Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Physical Adversarial Camouflage through Gradient Calibration and Regularization
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats
by: Liu, Kuanrong, et al.
Published: (2024)
by: Liu, Kuanrong, et al.
Published: (2024)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
by: Lou, Tianrui, et al.
Published: (2025)
by: Lou, Tianrui, et al.
Published: (2025)
Text Adversarial Attacks with Dynamic Outputs
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
Revisiting Backdoor Attacks against Large Vision-Language Models from Domain Shift
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives
by: Liu, Kuanrong, et al.
Published: (2025)
by: Liu, Kuanrong, et al.
Published: (2025)
Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Robust Anti-Backdoor Instruction Tuning in LVLMs
by: Xun, Yuan, et al.
Published: (2025)
by: Xun, Yuan, et al.
Published: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
BadCLIP: Dual-Embedding Guided Backdoor Attack on Multimodal Contrastive Learning
by: Liang, Siyuan, et al.
Published: (2023)
by: Liang, Siyuan, et al.
Published: (2023)
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
by: Liu, Enguang, et al.
Published: (2025)
by: Liu, Enguang, et al.
Published: (2025)
Multi-SIGATnet: A multimodal schizophrenia MRI classification algorithm using sparse interaction mechanisms and graph attention networks
by: Jiao, Yuhong, et al.
Published: (2024)
by: Jiao, Yuhong, et al.
Published: (2024)
Explaining latent representations of generative models with large multimodal models
by: Zhu, Mengdan, et al.
Published: (2024)
by: Zhu, Mengdan, et al.
Published: (2024)
A multimodal gesture recognition dataset for desktop human-computer interaction
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
by: Zhang, Runmin, et al.
Published: (2024)
by: Zhang, Runmin, et al.
Published: (2024)
SinSEMI: A One-Shot Image Generation Model and Data-Efficient Evaluation Framework for Semiconductor Inspection Equipment
by: Wu, ChunLiang, et al.
Published: (2025)
by: Wu, ChunLiang, et al.
Published: (2025)
FaceInsight: A Multimodal Large Language Model for Face Perception
by: Li, Jingzhi, et al.
Published: (2025)
by: Li, Jingzhi, et al.
Published: (2025)
UncTrack: Reliable Visual Object Tracking with Uncertainty-Aware Prototype Memory Network
by: Yao, Siyuan, et al.
Published: (2025)
by: Yao, Siyuan, et al.
Published: (2025)
Hide in Thicket: Generating Imperceptible and Rational Adversarial Perturbations on 3D Point Clouds
by: Lou, Tianrui, et al.
Published: (2024)
by: Lou, Tianrui, et al.
Published: (2024)
Enhancing multimodal cooperation via sample-level modality valuation
by: Wei, Yake, et al.
Published: (2023)
by: Wei, Yake, et al.
Published: (2023)
UAGLNet: Uncertainty-Aggregated Global-Local Fusion Network with Cooperative CNN-Transformer for Building Extraction
by: Yao, Siyuan, et al.
Published: (2025)
by: Yao, Siyuan, et al.
Published: (2025)
Mentor3AD: Feature Reconstruction-based 3D Anomaly Detection via Multi-modality Mentor Learning
by: Liang, Hanzhe
Published: (2025)
by: Liang, Hanzhe
Published: (2025)
From Image to Video, what do we need in multimodal LLMs?
by: Huang, Suyuan, et al.
Published: (2024)
by: Huang, Suyuan, et al.
Published: (2024)
Lie Detector: Unified Backdoor Detection via Cross-Examination Framework
by: Wang, Xuan, et al.
Published: (2025)
by: Wang, Xuan, et al.
Published: (2025)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
by: Yang, Junjie, et al.
Published: (2025)
by: Yang, Junjie, et al.
Published: (2025)
Hierarchical Graph Interaction Transformer with Dynamic Token Clustering for Camouflaged Object Detection
by: Yao, Siyuan, et al.
Published: (2024)
by: Yao, Siyuan, et al.
Published: (2024)
Similar Items
-
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025) -
Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
by: Chen, Yannan, et al.
Published: (2025) -
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
by: Chen, Ruoyu, et al.
Published: (2025) -
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024) -
Less is More: Fewer Interpretable Region via Submodular Subset Selection
by: Chen, Ruoyu, et al.
Published: (2024)