Beyond Visual Memory: Mechanistic Diagnostics of Latent Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Garvin, Chen, Yu, Wang, Xiang, Li, Shuai, Zhao, Xinpei, Liu, Huaxing, Dong, Shuai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
von: Guo, Garvin, et al.
Veröffentlicht: (2026)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
von: Wang, Qixun, et al.
Veröffentlicht: (2025)
Determined by User Needs: A Salient Object Detection Rationale Beyond Conventional Visual Stimuli
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
Dual Latent Memory for Visual Multi-agent System
von: Yu, Xinlei, et al.
Veröffentlicht: (2026)
von: Yu, Xinlei, et al.
Veröffentlicht: (2026)
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
von: Xi, Suyang, et al.
Veröffentlicht: (2026)
Beyond Static Visual Tokens: Structured Sequential Visual Chain-of-Thought Reasoning
von: Guo, Guangfu, et al.
Veröffentlicht: (2026)
von: Guo, Guangfu, et al.
Veröffentlicht: (2026)
LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
von: Hao, Haihong, et al.
Veröffentlicht: (2026)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoder
von: Liu, Tao, et al.
Veröffentlicht: (2023)
von: Liu, Tao, et al.
Veröffentlicht: (2023)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
von: Che, Liwei, et al.
Veröffentlicht: (2026)
von: Che, Liwei, et al.
Veröffentlicht: (2026)
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
von: Wu, Linquan, et al.
Veröffentlicht: (2026)
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
von: Zhang, Yang, et al.
Veröffentlicht: (2026)
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
von: Xue, Qiyao, et al.
Veröffentlicht: (2025)
von: Xue, Qiyao, et al.
Veröffentlicht: (2025)
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
von: Li, Zejun, et al.
Veröffentlicht: (2025)
von: Li, Zejun, et al.
Veröffentlicht: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
von: Li, Qiaoru, et al.
Veröffentlicht: (2026)
Beyond Pixels: Visual Metaphor Transfer via Schema-Driven Agentic Reasoning
von: Xu, Yu, et al.
Veröffentlicht: (2026)
von: Xu, Yu, et al.
Veröffentlicht: (2026)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
von: Shi, Fan, et al.
Veröffentlicht: (2025)
von: Shi, Fan, et al.
Veröffentlicht: (2025)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
von: Zhu, Chunzheng, et al.
Veröffentlicht: (2026)
von: Zhu, Chunzheng, et al.
Veröffentlicht: (2026)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
Exploring Visual Prompting: Robustness Inheritance and Beyond
von: Li, Qi, et al.
Veröffentlicht: (2025)
von: Li, Qi, et al.
Veröffentlicht: (2025)
Interleaved Latent Visual Reasoning with Selective Perceptual Modeling
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
von: Dong, Shuai, et al.
Veröffentlicht: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
Unbiased Visual Reasoning with Controlled Visual Inputs
von: Li, Zhaonan, et al.
Veröffentlicht: (2025)
von: Li, Zhaonan, et al.
Veröffentlicht: (2025)
SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual Datasets
von: Wan, Shenghua, et al.
Veröffentlicht: (2024)
von: Wan, Shenghua, et al.
Veröffentlicht: (2024)
Beyond Visual Safety: Jailbreaking Multimodal Large Language Models for Harmful Image Generation via Semantic-Agnostic Inputs
von: Yu, Mingyu, et al.
Veröffentlicht: (2026)
von: Yu, Mingyu, et al.
Veröffentlicht: (2026)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
von: Huang, Siyuan, et al.
Veröffentlicht: (2026)
CorrDetail: Visual Detail Enhanced Self-Correction for Face Forgery Detection
von: Zhou, Binjia, et al.
Veröffentlicht: (2025)
von: Zhou, Binjia, et al.
Veröffentlicht: (2025)
Predictive Reasoning with Augmented Anomaly Contrastive Learning for Compositional Visual Relations
von: Li, Chengtai, et al.
Veröffentlicht: (2026)
von: Li, Chengtai, et al.
Veröffentlicht: (2026)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
von: Jiang, Yanbei, et al.
Veröffentlicht: (2025)
von: Jiang, Yanbei, et al.
Veröffentlicht: (2025)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
von: Xie, Hongxia, et al.
Veröffentlicht: (2024)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attention
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
von: Tian, Changyuan, et al.
Veröffentlicht: (2026)
Enhancing Advanced Visual Reasoning Ability of Large Language Models
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2024)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
von: Dong, Zhuobai, et al.
Veröffentlicht: (2025)
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
von: Huang, Chi-Pin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
von: Guo, Garvin, et al.
Veröffentlicht: (2026) -
Monet: Reasoning in Latent Visual Space Beyond Images and Language
von: Wang, Qixun, et al.
Veröffentlicht: (2025) -
Determined by User Needs: A Salient Object Detection Rationale Beyond Conventional Visual Stimuli
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026) -
Dual Latent Memory for Visual Multi-agent System
von: Yu, Xinlei, et al.
Veröffentlicht: (2026) -
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
von: Xi, Suyang, et al.
Veröffentlicht: (2026)