Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Peng, Ruiying, Wu, Xueyu, Lei, Jing, Hou, Lu, Ma, Yuanzheng, Li, Xiaohui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
von: Ma, Ji, et al.
Veröffentlicht: (2026)
von: Ma, Ji, et al.
Veröffentlicht: (2026)
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
von: Peng, Ruiying, et al.
Veröffentlicht: (2026)
von: Peng, Ruiying, et al.
Veröffentlicht: (2026)
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
von: Zhang, Chi, et al.
Veröffentlicht: (2025)
Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
von: Zhang, Honglu, et al.
Veröffentlicht: (2025)
von: Zhang, Honglu, et al.
Veröffentlicht: (2025)
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
von: Lu, Wenting, et al.
Veröffentlicht: (2026)
von: Lu, Wenting, et al.
Veröffentlicht: (2026)
Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
von: Zuo, Jing, et al.
Veröffentlicht: (2026)
von: Zuo, Jing, et al.
Veröffentlicht: (2026)
Image-of-Thought Prompting for Visual Reasoning Refinement in Multimodal Large Language Models
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
von: Zhou, Qiji, et al.
Veröffentlicht: (2024)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
von: Li, Xiping, et al.
Veröffentlicht: (2025)
von: Li, Xiping, et al.
Veröffentlicht: (2025)
Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
Grounded Chain-of-Thought for Multimodal Large Language Models
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
von: Wu, Qiong, et al.
Veröffentlicht: (2025)
Exploring Perceptual Limitation of Multimodal Large Language Models
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
Multimodal Chain-of-Thought Reasoning in Language Models
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zhuosheng, et al.
Veröffentlicht: (2023)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shuoshuo, et al.
Veröffentlicht: (2025)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
von: Zhang, Jialiang, et al.
Veröffentlicht: (2026)
Attention-guided Fine-tuning of Multimodal Large Language Models Improves Chain-of-Thought Reasoning
von: Sinha, Sanchit, et al.
Veröffentlicht: (2026)
von: Sinha, Sanchit, et al.
Veröffentlicht: (2026)
Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
Understanding and Evaluating Hallucinations in 3D Visual Language Models
von: Peng, Ruiying, et al.
Veröffentlicht: (2025)
von: Peng, Ruiying, et al.
Veröffentlicht: (2025)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2025)
von: Liao, Yuan-Hong, et al.
Veröffentlicht: (2025)
Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
von: Lee, Jonathan, et al.
Veröffentlicht: (2025)
Brain-Inspired Capture: Evidence-Driven Neuromimetic Perceptual Simulation for Visual Decoding
von: Shao, Feixue, et al.
Veröffentlicht: (2026)
von: Shao, Feixue, et al.
Veröffentlicht: (2026)
Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
von: Zhong, Weihong, et al.
Veröffentlicht: (2024)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detection
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
von: Gu, Hexiang, et al.
Veröffentlicht: (2025)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
von: Cheng, Zihui, et al.
Veröffentlicht: (2025)
Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language Models
von: Zhang, Cheng, et al.
Veröffentlicht: (2026)
von: Zhang, Cheng, et al.
Veröffentlicht: (2026)
Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
von: Chen, Lei, et al.
Veröffentlicht: (2025)
von: Chen, Lei, et al.
Veröffentlicht: (2025)
Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks
von: Jing, Miao, et al.
Veröffentlicht: (2025)
von: Jing, Miao, et al.
Veröffentlicht: (2025)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
von: Huang, Muye, et al.
Veröffentlicht: (2025)
von: Huang, Muye, et al.
Veröffentlicht: (2025)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
von: He, Zhentao, et al.
Veröffentlicht: (2025)
von: He, Zhentao, et al.
Veröffentlicht: (2025)
Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA
von: Wu, Tong, et al.
Veröffentlicht: (2026)
von: Wu, Tong, et al.
Veröffentlicht: (2026)
SP-Mamba: Spatial-Perception State Space Model for Unsupervised Medical Anomaly Detection
von: Pan, Rui, et al.
Veröffentlicht: (2025)
von: Pan, Rui, et al.
Veröffentlicht: (2025)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Understanding and Mitigating Hallucinations in Multimodal Chain-of-Thought Models
von: Ma, Ji, et al.
Veröffentlicht: (2026) -
Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning
von: Peng, Ruiying, et al.
Veröffentlicht: (2026) -
Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
von: Zhang, Chi, et al.
Veröffentlicht: (2025) -
Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
von: Zhang, Honglu, et al.
Veröffentlicht: (2025) -
Watch Wider and Think Deeper: Collaborative Cross-modal Chain-of-Thought for Complex Visual Reasoning
von: Lu, Wenting, et al.
Veröffentlicht: (2026)