Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ou, Siqu, Wan, Tianrui, Zhao, Zhiyuan, Gao, Junyu, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
Exploring the Underwater World Segmentation without Extra Training
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
von: Li, Bingyu, et al.
Veröffentlicht: (2026)
MARIS: Marine Open-Vocabulary Instance Segmentation with Geometric Enhancement and Semantic Alignment
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
von: Li, Bingyu, et al.
Veröffentlicht: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives
von: Jiang, Kai, et al.
Veröffentlicht: (2025)
von: Jiang, Kai, et al.
Veröffentlicht: (2025)
What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language Models
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
von: Fan, Yingqi, et al.
Veröffentlicht: (2026)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
von: Kang, Seil, et al.
Veröffentlicht: (2025)
von: Kang, Seil, et al.
Veröffentlicht: (2025)
Do We Really Need a Large Number of Visual Prompts?
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
von: Kim, Youngeun, et al.
Veröffentlicht: (2023)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
von: Yue, Yang, et al.
Veröffentlicht: (2025)
von: Yue, Yang, et al.
Veröffentlicht: (2025)
Cross-Modal Projection in Multimodal LLMs Doesn't Really Project Visual Attributes to Textual Space
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
von: Verma, Gaurav, et al.
Veröffentlicht: (2024)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
von: Cao, Fanpu, et al.
Veröffentlicht: (2026)
von: Cao, Fanpu, et al.
Veröffentlicht: (2026)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2025)
Seeing Right but Saying Wrong: Inter- and Intra-Layer Refinement in MLLMs without Training
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
von: Song, Shezheng, et al.
Veröffentlicht: (2026)
How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images
von: Liu, Guimeng, et al.
Veröffentlicht: (2026)
von: Liu, Guimeng, et al.
Veröffentlicht: (2026)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
von: Ling, Chen, et al.
Veröffentlicht: (2026)
von: Ling, Chen, et al.
Veröffentlicht: (2026)
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
von: Wu, Xin, et al.
Veröffentlicht: (2026)
von: Wu, Xin, et al.
Veröffentlicht: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
U3M: Unbiased Multiscale Modal Fusion Model for Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
von: Li, Bingyu, et al.
Veröffentlicht: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
von: Liu, Zhining, et al.
Veröffentlicht: (2025)
von: Liu, Zhining, et al.
Veröffentlicht: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
GazeLLM: Multimodal LLMs incorporating Human Visual Attention
von: Rekimoto, Jun
Veröffentlicht: (2025)
von: Rekimoto, Jun
Veröffentlicht: (2025)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
von: Guo, Xingang, et al.
Veröffentlicht: (2025)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
von: Sun, Shichu, et al.
Veröffentlicht: (2025)
von: Sun, Shichu, et al.
Veröffentlicht: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaolin, et al.
Veröffentlicht: (2026)
Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2025)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2025)
CountQA: How Well Do MLLMs Count in the Wild?
von: Tamarapalli, Jayant Sravan, et al.
Veröffentlicht: (2025)
von: Tamarapalli, Jayant Sravan, et al.
Veröffentlicht: (2025)
Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
von: Li, Yuanshuai, et al.
Veröffentlicht: (2025)
Decoupling the Image Perception and Multimodal Reasoning for Reasoning Segmentation with Digital Twin Representations
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
von: Li, Yizhen, et al.
Veröffentlicht: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
von: Xiao, Tong, et al.
Veröffentlicht: (2025)
Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
von: Pang, Yuqi, et al.
Veröffentlicht: (2025)
Multimodal Fusion SLAM with Fourier Attention
von: Zhou, Youjie, et al.
Veröffentlicht: (2025)
von: Zhou, Youjie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2024) -
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025) -
Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
von: Li, Bingyu, et al.
Veröffentlicht: (2025) -
Exploring the Underwater World Segmentation without Extra Training
von: Li, Bingyu, et al.
Veröffentlicht: (2025) -
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
von: Li, Bingyu, et al.
Veröffentlicht: (2026)