Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ruoyu, Guo, Xiaoqing, Liu, Kangwei, Liang, Siyuan, Liu, Shiming, Zhang, Qunli, Wang, Laiyuan, Zhang, Hua, Cao, Xiaochun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making
by: Chen, Ruoyu, et al.
Published: (2026)
by: Chen, Ruoyu, et al.
Published: (2026)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
by: Chen, Yannan, et al.
Published: (2025)
by: Chen, Yannan, et al.
Published: (2025)
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Less is More: Fewer Interpretable Region via Submodular Subset Selection
by: Chen, Ruoyu, et al.
Published: (2024)
by: Chen, Ruoyu, et al.
Published: (2024)
Generalized Semantic Contrastive Learning via Embedding Side Information for Few-Shot Object Detection
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
by: Duan, Yuxiang, et al.
Published: (2025)
by: Duan, Yuxiang, et al.
Published: (2025)
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language Models
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
by: Fang, Haipeng, et al.
Published: (2025)
by: Fang, Haipeng, et al.
Published: (2025)
Efficient Backdoor Defense in Multimodal Contrastive Learning: A Token-Level Unlearning Method for Mitigating Threats
by: Liu, Kuanrong, et al.
Published: (2024)
by: Liu, Kuanrong, et al.
Published: (2024)
Bridging the Task Gap: Multi-Task Adversarial Transferability in CLIP and Its Derivatives
by: Liu, Kuanrong, et al.
Published: (2025)
by: Liu, Kuanrong, et al.
Published: (2025)
Unlearning Backdoor Threats: Enhancing Backdoor Defense in Multimodal Contrastive Learning via Local Token Unlearning
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
PhaseWin Search Framework Enable Efficient Object-Level Interpretation
by: Gu, Zihan, et al.
Published: (2025)
by: Gu, Zihan, et al.
Published: (2025)
Adversarial Backdoor Defense in CLIP
by: Kuang, Junhao, et al.
Published: (2024)
by: Kuang, Junhao, et al.
Published: (2024)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
by: Fu, Zixuan, et al.
Published: (2026)
by: Fu, Zixuan, et al.
Published: (2026)
GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
by: Zheng, Guanghao, et al.
Published: (2025)
by: Zheng, Guanghao, et al.
Published: (2025)
Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction
by: Guo, Ziyao, et al.
Published: (2025)
by: Guo, Ziyao, et al.
Published: (2025)
Object Detectors in the Open Environment: Challenges, Solutions, and Outlook
by: Liang, Siyuan, et al.
Published: (2024)
by: Liang, Siyuan, et al.
Published: (2024)
Hierarchical Graph Interaction Transformer with Dynamic Token Clustering for Camouflaged Object Detection
by: Yao, Siyuan, et al.
Published: (2024)
by: Yao, Siyuan, et al.
Published: (2024)
Continuous Speculative Decoding for Autoregressive Image Generation
by: Wang, Zili, et al.
Published: (2024)
by: Wang, Zili, et al.
Published: (2024)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
by: Zhang, Tao, et al.
Published: (2025)
by: Zhang, Tao, et al.
Published: (2025)
DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
by: Liu, Kangwei, et al.
Published: (2025)
by: Liu, Kangwei, et al.
Published: (2025)
From "What" to "How": Constrained Reasoning for Autoregressive Image Generation
by: Yan, Ruxue, et al.
Published: (2026)
by: Yan, Ruxue, et al.
Published: (2026)
3D Gaussian Splatting Driven Multi-View Robust Physical Adversarial Camouflage Generation
by: Lou, Tianrui, et al.
Published: (2025)
by: Lou, Tianrui, et al.
Published: (2025)
Text Adversarial Attacks with Dynamic Outputs
by: Wang, Wenqiang, et al.
Published: (2025)
by: Wang, Wenqiang, et al.
Published: (2025)
R-PGA: Robust Physical Adversarial Camouflage Generation via Relightable 3D Gaussian Splatting
by: Lou, Tianrui, et al.
Published: (2026)
by: Lou, Tianrui, et al.
Published: (2026)
Poisoned Forgery Face: Towards Backdoor Attacks on Face Forgery Detection
by: Liang, Jiawei, et al.
Published: (2024)
by: Liang, Jiawei, et al.
Published: (2024)
Exploring Inconsistent Knowledge Distillation for Object Detection with Data Augmentation
by: Liang, Jiawei, et al.
Published: (2022)
by: Liang, Jiawei, et al.
Published: (2022)
Physical Adversarial Camouflage through Gradient Calibration and Regularization
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
Robust Anti-Backdoor Instruction Tuning in LVLMs
by: Xun, Yuan, et al.
Published: (2025)
by: Xun, Yuan, et al.
Published: (2025)
CleanerCLIP: Fine-grained Counterfactual Semantic Augmentation for Backdoor Defense in Contrastive Learning
by: Xun, Yuan, et al.
Published: (2024)
by: Xun, Yuan, et al.
Published: (2024)
FaceInsight: A Multimodal Large Language Model for Face Perception
by: Li, Jingzhi, et al.
Published: (2025)
by: Li, Jingzhi, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
SuperVoxelGPT: Adaptive and Ordered 3D Tokenization for Autoregressive Shape Generation
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
TokenAR: Multiple Subject Generation via Autoregressive Token-level enhancement
by: Sun, Haiyue, et al.
Published: (2025)
by: Sun, Haiyue, et al.
Published: (2025)
From Visuals to Vocabulary: Establishing Equivalence Between Image and Text Token Through Autoregressive Pre-training in MLLMs
by: Li, Mingxiao, et al.
Published: (2025)
by: Li, Mingxiao, et al.
Published: (2025)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
Similar Items
-
Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making
by: Chen, Ruoyu, et al.
Published: (2026) -
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025) -
Did Models Sufficient Learn? Attribution-Guided Training via Subset-Selected Counterfactual Augmentation
by: Chen, Yannan, et al.
Published: (2025) -
Less is More: Efficient Black-box Attribution via Minimal Interpretable Subset Selection
by: Chen, Ruoyu, et al.
Published: (2025) -
Interpreting Object-level Foundation Models via Visual Precision Search
by: Chen, Ruoyu, et al.
Published: (2024)