Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Qiyan, Zhang, Xiaofeng, Chang, Shuochen, Chen, Qianyu, Yuan, Xiaosong, Chen, Xuhang, Liu, Luoqi, Zhang, Jiajun, Zhang, Xu-Yao, Wang, Da-Han |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion-CAM: Faithful Visual Explanations for dMLLMs
by: Zuo, Haomin, et al.
Published: (2026)
by: Zuo, Haomin, et al.
Published: (2026)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
by: Chang, Shuochen, et al.
Published: (2025)
by: Chang, Shuochen, et al.
Published: (2025)
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025)
by: Zhao, Qiyan, et al.
Published: (2025)
Understanding In-Context Learning from Repetitions
by: Yan, Jianhao, et al.
Published: (2023)
by: Yan, Jianhao, et al.
Published: (2023)
Image Generation Based on Image Style Extraction
by: Chang, Shuochen
Published: (2025)
by: Chang, Shuochen
Published: (2025)
Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Do MLLMs Really Understand the Charts?
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Understanding the RoPE Extensions of Long-Context LLMs: An Attention Perspective
by: Zhong, Meizhi, et al.
Published: (2024)
by: Zhong, Meizhi, et al.
Published: (2024)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
by: Zhang, Qizhe, et al.
Published: (2025)
by: Zhang, Qizhe, et al.
Published: (2025)
Understanding the Repeat Curse in Large Language Models from a Feature Perspective
by: Yao, Junchi, et al.
Published: (2025)
by: Yao, Junchi, et al.
Published: (2025)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
by: Xi, Yingjie, et al.
Published: (2025)
by: Xi, Yingjie, et al.
Published: (2025)
ART: Attention Replacement Technique to Improve Factuality in LLMs
by: Luo, Ziqin, et al.
Published: (2026)
by: Luo, Ziqin, et al.
Published: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
by: Zhang, Junyang, et al.
Published: (2025)
by: Zhang, Junyang, et al.
Published: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
by: Zhang, Gengyuan, et al.
Published: (2025)
by: Zhang, Gengyuan, et al.
Published: (2025)
Memory Efficient Matting with Adaptive Token Routing
by: Lin, Yiheng, et al.
Published: (2024)
by: Lin, Yiheng, et al.
Published: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
by: Zou, Xin, et al.
Published: (2025)
by: Zou, Xin, et al.
Published: (2025)
Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
by: Chen, Ruoyu, et al.
Published: (2025)
by: Chen, Ruoyu, et al.
Published: (2025)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
by: Qu, Liao, et al.
Published: (2024)
by: Qu, Liao, et al.
Published: (2024)
Hallucination Begins Where Saliency Drops
by: Zhang, Xiaofeng, et al.
Published: (2026)
by: Zhang, Xiaofeng, et al.
Published: (2026)
Reasoning Fails Where Step Flow Breaks
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
by: Zhang, Chao, et al.
Published: (2025)
by: Zhang, Chao, et al.
Published: (2025)
Understanding the Curse of Unrolling
by: Mehmood, Sheheryar, et al.
Published: (2026)
by: Mehmood, Sheheryar, et al.
Published: (2026)
AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
by: Kitouni, Ouail, et al.
Published: (2024)
by: Kitouni, Ouail, et al.
Published: (2024)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
by: Liu, Can, et al.
Published: (2025)
by: Liu, Can, et al.
Published: (2025)
TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
by: Zhang, Kejia, et al.
Published: (2025)
by: Zhang, Kejia, et al.
Published: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
by: Sun, Yanpeng, et al.
Published: (2025)
by: Sun, Yanpeng, et al.
Published: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
MP-GUI: Modality Perception with MLLMs for GUI Understanding
by: Wang, Ziwei, et al.
Published: (2025)
by: Wang, Ziwei, et al.
Published: (2025)
Understanding and Controlling Repetition Neurons and Induction Heads in In-Context Learning
by: Doan, Nhi Hoai, et al.
Published: (2025)
by: Doan, Nhi Hoai, et al.
Published: (2025)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
Visual Room 2.0: Seeing is Not Understanding for MLLMs
by: Li, Haokun, et al.
Published: (2025)
by: Li, Haokun, et al.
Published: (2025)
Contrastive Token Learning with Similarity Decay for Repetition Suppression in Machine Translation
by: Dai, Huangyu, et al.
Published: (2024)
by: Dai, Huangyu, et al.
Published: (2024)
Empowering Functional Neuroimaging: A Pre-trained Generative Framework for Unified Representation of Neural Signals
by: Yao, Weiheng, et al.
Published: (2025)
by: Yao, Weiheng, et al.
Published: (2025)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
by: Yue, Zhengrong, et al.
Published: (2025)
by: Yue, Zhengrong, et al.
Published: (2025)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
MSSP : A Versatile Multi-Scenario Adaptable Intelligent Robot Simulation Platform Based on LIDAR-Inertial Fusion
by: Li, Qiyan, et al.
Published: (2024)
by: Li, Qiyan, et al.
Published: (2024)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
by: Chen, Feng, et al.
Published: (2025)
by: Chen, Feng, et al.
Published: (2025)
Similar Items
-
Diffusion-CAM: Faithful Visual Explanations for dMLLMs
by: Zuo, Haomin, et al.
Published: (2026) -
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
by: Chang, Shuochen, et al.
Published: (2025) -
MCA-LLaVA: Manhattan Causal Attention for Reducing Hallucination in Large Vision-Language Models
by: Zhao, Qiyan, et al.
Published: (2025) -
Understanding In-Context Learning from Repetitions
by: Yan, Jianhao, et al.
Published: (2023) -
Image Generation Based on Image Style Extraction
by: Chang, Shuochen
Published: (2025)