Language Models Can See Better: Visual Contrastive Decoding For LLM Multimodal Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Yuqi, Yang, Bowen, Tu, Haoqin, Cao, Yun, Zhang, Zeyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
by: He, Yiran, et al.
Published: (2025)
by: He, Yiran, et al.
Published: (2025)
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025)
by: Pang, Yuqi, et al.
Published: (2025)
Multimodal Language Models See Better When They Look Shallower
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
BLINK: Multimodal Large Language Models Can See but Not Perceive
by: Fu, Xingyu, et al.
Published: (2024)
by: Fu, Xingyu, et al.
Published: (2024)
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)
by: Park, Yeji, et al.
Published: (2024)
Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
by: Song, Python, et al.
Published: (2025)
by: Song, Python, et al.
Published: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
by: Upadhyay, Ujjwal, et al.
Published: (2025)
by: Upadhyay, Ujjwal, et al.
Published: (2025)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
by: Yang, Bowen, et al.
Published: (2025)
by: Yang, Bowen, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
by: Huang, Zheng, et al.
Published: (2025)
by: Huang, Zheng, et al.
Published: (2025)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
by: Lee, Yi-Lun, et al.
Published: (2024)
by: Lee, Yi-Lun, et al.
Published: (2024)
LANTERN: Accelerating Visual Autoregressive Models with Relaxed Speculative Decoding
by: Jang, Doohyuk, et al.
Published: (2024)
by: Jang, Doohyuk, et al.
Published: (2024)
Integrating Multimodal Large Language Model Knowledge into Amodal Completion
by: Yun, Heecheol, et al.
Published: (2026)
by: Yun, Heecheol, et al.
Published: (2026)
The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
by: Goyal, Karan
Published: (2026)
by: Goyal, Karan
Published: (2026)
Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models
by: Zhang, Jiaming, et al.
Published: (2025)
by: Zhang, Jiaming, et al.
Published: (2025)
Speculative Decoding Reimagined for Multimodal Large Language Models
by: Lin, Luxi, et al.
Published: (2025)
by: Lin, Luxi, et al.
Published: (2025)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
Seeing the Image: Prioritizing Visual Correlation by Contrastive Alignment
by: Xiao, Xin, et al.
Published: (2024)
by: Xiao, Xin, et al.
Published: (2024)
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
by: Li, Pengteng, et al.
Published: (2025)
by: Li, Pengteng, et al.
Published: (2025)
Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
by: Fu, Rao, et al.
Published: (2024)
by: Fu, Rao, et al.
Published: (2024)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
by: Wang, Yuqing, et al.
Published: (2023)
by: Wang, Yuqing, et al.
Published: (2023)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models
by: Lee, Jihoon, et al.
Published: (2025)
by: Lee, Jihoon, et al.
Published: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
by: Lei, Zeyu, et al.
Published: (2025)
by: Lei, Zeyu, et al.
Published: (2025)
Can NeRFs See without Cameras?
by: Amballa, Chaitanya, et al.
Published: (2025)
by: Amballa, Chaitanya, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
by: Shen, Yixuan, et al.
Published: (2026)
by: Shen, Yixuan, et al.
Published: (2026)
Reasoning-Aware Multimodal Fusion for Hateful Video Detection
by: Yang, Shuonan, et al.
Published: (2025)
by: Yang, Shuonan, et al.
Published: (2025)
A Survey on Visual Anomaly Detection: Challenge, Approach, and Prospect
by: Cao, Yunkang, et al.
Published: (2024)
by: Cao, Yunkang, et al.
Published: (2024)
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
by: Yan, Hao, et al.
Published: (2025)
by: Yan, Hao, et al.
Published: (2025)
Improve Vision Language Model Chain-of-thought Reasoning
by: Zhang, Ruohong, et al.
Published: (2024)
by: Zhang, Ruohong, et al.
Published: (2024)
Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Prompting
by: Frisoni, Giacomo, et al.
Published: (2026)
by: Frisoni, Giacomo, et al.
Published: (2026)
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
by: Tang, Jianting, et al.
Published: (2025)
by: Tang, Jianting, et al.
Published: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
by: Zheng, Haojie, et al.
Published: (2024)
by: Zheng, Haojie, et al.
Published: (2024)
Similar Items
-
Can GPT tell us why these images are synthesized? Empowering Multimodal Large Language Models for Forensics
by: He, Yiran, et al.
Published: (2025) -
MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
by: Pang, Yuqi, et al.
Published: (2025) -
Multimodal Language Models See Better When They Look Shallower
by: Chen, Haoran, et al.
Published: (2025) -
BLINK: Multimodal Large Language Models Can See but Not Perceive
by: Fu, Xingyu, et al.
Published: (2024) -
ConVis: Contrastive Decoding with Hallucination Visualization for Mitigating Hallucinations in Multimodal Large Language Models
by: Park, Yeji, et al.
Published: (2024)