Bad Seeing or Bad Thinking? Rewarding Perception for Vision-Language Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Haozhe, Xu, Qixin, Wang, Changpeng, Xue, Taofeng, Peng, Chong, Chen, Wenhu, Lin, Fangzhen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
by: Wang, Changpeng, et al.
Published: (2025)
by: Wang, Changpeng, et al.
Published: (2025)
CogDoc: Towards Unified thinking in Documents
by: Xu, Qixin, et al.
Published: (2025)
by: Xu, Qixin, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
by: Wu, Yuhuan, et al.
Published: (2026)
by: Wu, Yuhuan, et al.
Published: (2026)
BadHMP: Backdoor Attack against Human Motion Prediction
by: Xu, Chaohui, et al.
Published: (2024)
by: Xu, Chaohui, et al.
Published: (2024)
BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
by: Wang, Ruotong, et al.
Published: (2025)
by: Wang, Ruotong, et al.
Published: (2025)
BadViM: Backdoor Attack against Vision Mamba
by: Wu, Yinghao, et al.
Published: (2025)
by: Wu, Yinghao, et al.
Published: (2025)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Reward Guided Latent Consistency Distillation
by: Li, Jiachen, et al.
Published: (2024)
by: Li, Jiachen, et al.
Published: (2024)
Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation
by: Zhu, Xiaomeng, et al.
Published: (2025)
by: Zhu, Xiaomeng, et al.
Published: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Are Hallucinations Bad Estimations?
by: Liu, Hude, et al.
Published: (2025)
by: Liu, Hude, et al.
Published: (2025)
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
by: Wu, Juncheng, et al.
Published: (2026)
by: Wu, Juncheng, et al.
Published: (2026)
Block-wise LoRA: Revisiting Fine-grained LoRA for Effective Personalization and Stylization in Text-to-Image Generation
by: Li, Likun, et al.
Published: (2024)
by: Li, Likun, et al.
Published: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Information-Theoretic Constraints for Continual Vision-Language-Action Alignment
by: Zhao, Libang, et al.
Published: (2026)
by: Zhao, Libang, et al.
Published: (2026)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
by: Zhang, Jialiang, et al.
Published: (2026)
by: Zhang, Jialiang, et al.
Published: (2026)
BadSR: Stealthy Label Backdoor Attacks on Image Super-Resolution
by: Guo, Ji, et al.
Published: (2025)
by: Guo, Ji, et al.
Published: (2025)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
by: Wu, Zhiheng, et al.
Published: (2026)
by: Wu, Zhiheng, et al.
Published: (2026)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
by: Wu, Junfei, et al.
Published: (2025)
by: Wu, Junfei, et al.
Published: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
by: Liu, Chengzhi, et al.
Published: (2025)
by: Liu, Chengzhi, et al.
Published: (2025)
Seeing the Abstract: Translating the Abstract Language for Vision Language Models
by: Talon, Davide, et al.
Published: (2025)
by: Talon, Davide, et al.
Published: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
by: Wang, Chaoyang, et al.
Published: (2026)
by: Wang, Chaoyang, et al.
Published: (2026)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
by: Wang, Haibo, et al.
Published: (2026)
by: Wang, Haibo, et al.
Published: (2026)
Do Vision-Language Models See Urban Scenes as People Do? An Urban Perception Benchmark
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
Good Scores, Bad Data: A Metric for Multimodal Coherence
by: Srinivasan, Vasundra
Published: (2026)
by: Srinivasan, Vasundra
Published: (2026)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
TINA: Think, Interaction, and Action Framework for Zero-Shot Vision Language Navigation
by: Li, Dingbang, et al.
Published: (2024)
by: Li, Dingbang, et al.
Published: (2024)
GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
by: Guo, Shasha, et al.
Published: (2025)
by: Guo, Shasha, et al.
Published: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
MediSee: Reasoning-based Pixel-level Perception in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
Reward Design for Physical Reasoning in Vision-Language Models
by: Lilienthal, Derek, et al.
Published: (2026)
by: Lilienthal, Derek, et al.
Published: (2026)
Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
by: Guo, Xingang, et al.
Published: (2025)
by: Guo, Xingang, et al.
Published: (2025)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
by: Xu, Haolei, et al.
Published: (2026)
by: Xu, Haolei, et al.
Published: (2026)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
BadPatch: Diffusion-Based Generation of Physical Adversarial Patches
by: Wang, Zhixiang, et al.
Published: (2024)
by: Wang, Zhixiang, et al.
Published: (2024)
Similar Items
-
From Illusion to Intention: Visual Rationale Learning for Vision-Language Reasoning
by: Wang, Changpeng, et al.
Published: (2025) -
CogDoc: Towards Unified thinking in Documents
by: Xu, Qixin, et al.
Published: (2025) -
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025) -
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
by: Wu, Yuhuan, et al.
Published: (2026) -
BadHMP: Backdoor Attack against Human Motion Prediction
by: Xu, Chaohui, et al.
Published: (2024)