Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Xin, Lu, Di, Wang, Yizhou, Yan, Yibo, Lyu, Yuanhuiyi, Zheng, Xu, Zhang, Linfeng, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024)
by: Li, Jungang, et al.
Published: (2024)
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025)
by: Huo, Jiahao, et al.
Published: (2025)
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)
by: Zou, Xin, et al.
Published: (2024)
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
Don't Hesitate, Just Collaborate!
by: Burk, Lynne F.
Published: (2007)
by: Burk, Lynne F.
Published: (2007)
Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
by: Lyu, Yuanhuiyi, et al.
Published: (2025)
Don't Just Fine-tune the Agent, Tune the Environment
by: Lu, Siyuan, et al.
Published: (2025)
by: Lu, Siyuan, et al.
Published: (2025)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
by: Zheng, Kening, et al.
Published: (2024)
by: Zheng, Kening, et al.
Published: (2024)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
by: Zhong, Ding, et al.
Published: (2025)
by: Zhong, Ding, et al.
Published: (2025)
Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event Camera
by: Cao, Jiahang, et al.
Published: (2023)
by: Cao, Jiahang, et al.
Published: (2023)
From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Unlocking Multimodal Document Intelligence: From Current Triumphs to Future Frontiers of Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Position: Measure Dataset Diversity, Don't Just Claim It
by: Zhao, Dora, et al.
Published: (2024)
by: Zhao, Dora, et al.
Published: (2024)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
by: Yan, Yibo, et al.
Published: (2025)
by: Yan, Yibo, et al.
Published: (2025)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
by: Deng, Haolin, et al.
Published: (2026)
by: Deng, Haolin, et al.
Published: (2026)
MLLMs are Deeply Affected by Modality Bias
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering
by: Liu, Shuliang, et al.
Published: (2026)
by: Liu, Shuliang, et al.
Published: (2026)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
by: Dang, Yunkai, et al.
Published: (2024)
by: Dang, Yunkai, et al.
Published: (2024)
GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning
by: Zhang, Jianghangfan, et al.
Published: (2025)
by: Zhang, Jianghangfan, et al.
Published: (2025)
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges
by: Yan, Yibo, et al.
Published: (2024)
by: Yan, Yibo, et al.
Published: (2024)
Visual Late Chunking: An Empirical Study of Contextual Chunking for Efficient Visual Document Retrieval
by: Yan, Yibo, et al.
Published: (2026)
by: Yan, Yibo, et al.
Published: (2026)
Don't Be Greedy, Just Relax! Pruning LLMs via Frank-Wolfe
by: Roux, Christophe, et al.
Published: (2025)
by: Roux, Christophe, et al.
Published: (2025)
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
by: Cao, Mingyu, et al.
Published: (2024)
by: Cao, Mingyu, et al.
Published: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarification
by: Yang, Yuting, et al.
Published: (2026)
by: Yang, Yuting, et al.
Published: (2026)
Unlocking Speech Instruction Data Potential with Query Rewriting
by: Hei, Yonghua, et al.
Published: (2025)
by: Hei, Yonghua, et al.
Published: (2025)
'Show It, Don't Just Say It': The Complementary Effects of Instruction Multimodality for Software Guidance
by: Poh, Emran, et al.
Published: (2026)
by: Poh, Emran, et al.
Published: (2026)
We Don't Have to Learn Anything; We Just Have to Find the Answer
by: Hoover, Clara
Published: (2005)
by: Hoover, Clara
Published: (2005)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
Mitigating Modality Prior-Induced Hallucinations in Multimodal Large Language Models via Deciphering Attention Causality
by: Zhou, Guanyu, et al.
Published: (2024)
by: Zhou, Guanyu, et al.
Published: (2024)
Decoding by Perturbation: Mitigating MLLM Hallucinations via Dynamic Textual Perturbation
by: Jia, Sihang, et al.
Published: (2026)
by: Jia, Sihang, et al.
Published: (2026)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Similar Items
-
SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
by: Li, Jungang, et al.
Published: (2024) -
RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive Learning
by: Lyu, Yuanhuiyi, et al.
Published: (2025) -
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025) -
MMUnlearner: Reformulating Multimodal Machine Unlearning in the Era of Multimodal Large Language Models
by: Huo, Jiahao, et al.
Published: (2025) -
Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models
by: Zou, Xin, et al.
Published: (2024)