Diagnosing Visual Reasoning: Challenges, Insights, and a Path Forward
Fuente:
arXiv
Saved in:
| Main Authors: | Bi, Jing, Sun, Guangyu, Vosoughi, Ali, Chen, Chen, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
EAGLE: Egocentric AGgregated Language-video Engine
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
What to Do Next? Memorizing skills from Egocentric Instructional Video
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
by: Chia, Yew Ken, et al.
Published: (2024)
by: Chia, Yew Ken, et al.
Published: (2024)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
by: Huang, Yixu, et al.
Published: (2026)
by: Huang, Yixu, et al.
Published: (2026)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
by: Zhang, Ye, et al.
Published: (2025)
by: Zhang, Ye, et al.
Published: (2025)
Semantic-Enriched Latent Visual Reasoning
by: Xu, Tianrun, et al.
Published: (2026)
by: Xu, Tianrun, et al.
Published: (2026)
FreSca: Scaling in Frequency Space Enhances Diffusion Models
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
by: Hong, Hyesoo, et al.
Published: (2026)
by: Hong, Hyesoo, et al.
Published: (2026)
The Role of Visual Modality in Multimodal Mathematical Reasoning: Challenges and Insights
by: Liu, Yufang, et al.
Published: (2025)
by: Liu, Yufang, et al.
Published: (2025)
Forward Learning for Gradient-based Black-box Saliency Map Generation
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation
by: Liang, Susan, et al.
Published: (2024)
by: Liang, Susan, et al.
Published: (2024)
Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
by: Fan, Zhiyuan, et al.
Published: (2025)
by: Fan, Zhiyuan, et al.
Published: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics
by: Lan, Jian, et al.
Published: (2026)
by: Lan, Jian, et al.
Published: (2026)
Latent Visual Reasoning
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
TraversalBench: Challenging Paths to Follow for Vision Language Models
by: Petrova, Clara, et al.
Published: (2026)
by: Petrova, Clara, et al.
Published: (2026)
Benchmarking Object Detectors with COCO: A New Path Forward
by: Singh, Shweta, et al.
Published: (2024)
by: Singh, Shweta, et al.
Published: (2024)
AnchorSplat: Feed-Forward 3D Gaussian Splatting with 3D Geometric Priors
by: Zhang, Xiaoxue, et al.
Published: (2026)
by: Zhang, Xiaoxue, et al.
Published: (2026)
EGGS: Exchangeable 2D/3D Gaussian Splatting for Geometry-Appearance Balanced Novel View Synthesis
by: Zhang, Yancheng, et al.
Published: (2025)
by: Zhang, Yancheng, et al.
Published: (2025)
Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Tri$^{2}$-plane: Thinking Head Avatar via Feature Pyramid
by: Song, Luchuan, et al.
Published: (2024)
by: Song, Luchuan, et al.
Published: (2024)
Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
by: Dong, Yuhao, et al.
Published: (2024)
by: Dong, Yuhao, et al.
Published: (2024)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
by: Tang, Yunlong, et al.
Published: (2025)
by: Tang, Yunlong, et al.
Published: (2025)
Efficiently Leveraging Linguistic Priors for Scene Text Spotting
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing
by: Khalid, Umar, et al.
Published: (2024)
by: Khalid, Umar, et al.
Published: (2024)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum
by: Zhang, Yuliang, et al.
Published: (2026)
by: Zhang, Yuliang, et al.
Published: (2026)
Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge
by: Xu, Yinsong, et al.
Published: (2026)
by: Xu, Yinsong, et al.
Published: (2026)
HCL-FF: Hierarchical and Contrastive Learning for Forward-Forward Algorithm
by: Yao, Jie-En, et al.
Published: (2026)
by: Yao, Jie-En, et al.
Published: (2026)
Expedited Training of Visual Conditioned Language Generation via Redundancy Reduction
by: Jian, Yiren, et al.
Published: (2023)
by: Jian, Yiren, et al.
Published: (2023)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Similar Items
-
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity
by: Bi, Jing, et al.
Published: (2025) -
EAGLE: Egocentric AGgregated Language-video Engine
by: Bi, Jing, et al.
Published: (2024) -
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024) -
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025) -
What to Do Next? Memorizing skills from Egocentric Instructional Video
by: Bi, Jing, et al.
Published: (2025)