More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Xinyu, Zou, Shu, Yang, Zhaoyuan, He, Mengqi, Waschkowski, Fabian, Wesemann, Lukas, Tu, Peter, Zhang, Jing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
by: Zou, Shu, et al.
Published: (2025)
by: Zou, Shu, et al.
Published: (2025)
All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2026)
by: Tian, Xinyu, et al.
Published: (2026)
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2023)
by: Tian, Xinyu, et al.
Published: (2023)
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)
by: He, Mengqi, et al.
Published: (2025)
Black Sheep in the Herd: Playing with Spuriously Correlated Attributes for Vision-Language Recognition
by: Tian, Xinyu, et al.
Published: (2025)
by: Tian, Xinyu, et al.
Published: (2025)
SimLabel: Consistency-Guided OOD Detection with Pretrained Vision-Language Models
by: Zou, Shu, et al.
Published: (2025)
by: Zou, Shu, et al.
Published: (2025)
Robust Spatiotemporal Forecasting Using Adaptive Deep-Unfolded Variational Mode Decomposition
by: Ahmad, Osama, et al.
Published: (2025)
by: Ahmad, Osama, et al.
Published: (2025)
Variational Mode-Driven Graph Convolutional Network for Spatiotemporal Traffic Forecasting
by: Ahmad, Osama, et al.
Published: (2024)
by: Ahmad, Osama, et al.
Published: (2024)
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
by: He, Mengqi, et al.
Published: (2026)
by: He, Mengqi, et al.
Published: (2026)
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
by: Wu, Shaojin, et al.
Published: (2025)
by: Wu, Shaojin, et al.
Published: (2025)
Less is More: Lean yet Powerful Vision-Language Model for Autonomous Driving
by: Yang, Sheng, et al.
Published: (2025)
by: Yang, Sheng, et al.
Published: (2025)
DualCoT-VLA: Visual-Linguistic Chain of Thought via Parallel Reasoning for Vision-Language-Action Models
by: Zhong, Zhide, et al.
Published: (2026)
by: Zhong, Zhide, et al.
Published: (2026)
FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation
by: Zuo, Jing, et al.
Published: (2026)
by: Zuo, Jing, et al.
Published: (2026)
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
by: Xie, Qinghongbing, et al.
Published: (2025)
by: Xie, Qinghongbing, et al.
Published: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
by: Li, Xuchen, et al.
Published: (2025)
by: Li, Xuchen, et al.
Published: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
by: Chen, Yangyi, et al.
Published: (2023)
by: Chen, Yangyi, et al.
Published: (2023)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
DreamSteerer: Enhancing Source Image Conditioned Editability using Personalized Diffusion Models
by: Yu, Zhengyang, et al.
Published: (2024)
by: Yu, Zhengyang, et al.
Published: (2024)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
by: Pham, Tan-Hanh, et al.
Published: (2025)
by: Pham, Tan-Hanh, et al.
Published: (2025)
C-CoT: Counterfactual Chain-of-Thought with Vision-Language Models for Safe Autonomous Driving
by: Tian, Kefei, et al.
Published: (2026)
by: Tian, Kefei, et al.
Published: (2026)
NuScenes-SpatialQA: A Spatial Understanding and Reasoning Benchmark for Vision-Language Models in Autonomous Driving
by: Tian, Kexin, et al.
Published: (2025)
by: Tian, Kexin, et al.
Published: (2025)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
by: Wang, Wenchuan, et al.
Published: (2025)
by: Wang, Wenchuan, et al.
Published: (2025)
Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
by: Bouguerra, Aymen, et al.
Published: (2025)
by: Bouguerra, Aymen, et al.
Published: (2025)
LPT: Less-overfitting Prompt Tuning for Vision-Language Model
by: Ding, Chenhao, et al.
Published: (2024)
by: Ding, Chenhao, et al.
Published: (2024)
AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning
by: Li, Xiping, et al.
Published: (2025)
by: Li, Xiping, et al.
Published: (2025)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
by: Zhang, Jialiang, et al.
Published: (2026)
by: Zhang, Jialiang, et al.
Published: (2026)
IMPUS: Image Morphing with Perceptually-Uniform Sampling Using Diffusion Models
by: Yang, Zhaoyuan, et al.
Published: (2023)
by: Yang, Zhaoyuan, et al.
Published: (2023)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
by: Man, Yunze, et al.
Published: (2025)
by: Man, Yunze, et al.
Published: (2025)
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
by: Li, Tianqin, et al.
Published: (2025)
by: Li, Tianqin, et al.
Published: (2025)
Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
TinySAM 2: Extreme Memory Compression for Efficient Track Anything Model
by: Ding, Zhaoyuan, et al.
Published: (2026)
by: Ding, Zhaoyuan, et al.
Published: (2026)
The Dual Mechanisms of Spatial Reasoning in Vision-Language Models
by: Cui, Kelly, et al.
Published: (2026)
by: Cui, Kelly, et al.
Published: (2026)
Less is More: Discovering Concise Network Explanations
by: Kondapaneni, Neehar, et al.
Published: (2024)
by: Kondapaneni, Neehar, et al.
Published: (2024)
Probability Density Geodesics in Image Diffusion Latent Space
by: Yu, Qingtao, et al.
Published: (2025)
by: Yu, Qingtao, et al.
Published: (2025)
AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving
by: Qian, Kangan, et al.
Published: (2025)
by: Qian, Kangan, et al.
Published: (2025)
Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning
by: Deng, Huilin, et al.
Published: (2025)
by: Deng, Huilin, et al.
Published: (2025)
Similar Items
-
Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
by: Zou, Shu, et al.
Published: (2025) -
All Roads Lead to Rome: Incentivizing Divergent Thinking in Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2026) -
Identifying and Mitigating Position Bias of Multi-image Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2025) -
ArGue: Attribute-Guided Prompt Tuning for Vision-Language Models
by: Tian, Xinyu, et al.
Published: (2023) -
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
by: He, Mengqi, et al.
Published: (2025)