Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | He, Jixuan, Li, Xueting, Lin, Chieh Hubert, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
by: He, Jixuan, et al.
Published: (2025)
by: He, Jixuan, et al.
Published: (2025)
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
by: You, Junqi, et al.
Published: (2025)
by: You, Junqi, et al.
Published: (2025)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Gaga: Group Any Gaussians via 3D-aware Memory Bank
by: Lyu, Weijie, et al.
Published: (2024)
by: Lyu, Weijie, et al.
Published: (2024)
Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint
by: Zhou, Junwei, et al.
Published: (2024)
by: Zhou, Junwei, et al.
Published: (2024)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
by: Bao, Jingzhi, et al.
Published: (2024)
by: Bao, Jingzhi, et al.
Published: (2024)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
CoCo4D: Comprehensive and Complex 4D Scene Generation
by: Zhou, Junwei, et al.
Published: (2025)
by: Zhou, Junwei, et al.
Published: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
by: Wang, Junyang, et al.
Published: (2026)
by: Wang, Junyang, et al.
Published: (2026)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
by: Liu, Jiageng, et al.
Published: (2025)
by: Liu, Jiageng, et al.
Published: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
by: Qiao, Yuxuan, et al.
Published: (2024)
by: Qiao, Yuxuan, et al.
Published: (2024)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
On the Perception Bottleneck of VLMs for Chart Understanding
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
by: Avogaro, Niccolo, et al.
Published: (2026)
by: Avogaro, Niccolo, et al.
Published: (2026)
Synthesizing Consistent Novel Views via 3D Epipolar Attention without Re-Training
by: Ye, Botao, et al.
Published: (2025)
by: Ye, Botao, et al.
Published: (2025)
Pyramid Diffusion for Fine 3D Large Scene Generation
by: Liu, Yuheng, et al.
Published: (2023)
by: Liu, Yuheng, et al.
Published: (2023)
3D Reconstruction with Spatial Memory
by: Wang, Hengyi, et al.
Published: (2024)
by: Wang, Hengyi, et al.
Published: (2024)
Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
by: Lu, Meng, et al.
Published: (2025)
by: Lu, Meng, et al.
Published: (2025)
Are VLMs Really Blind
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
by: Zhan, Shaoxiong, et al.
Published: (2026)
by: Zhan, Shaoxiong, et al.
Published: (2026)
Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
by: Wang, Dianyi, et al.
Published: (2025)
by: Wang, Dianyi, et al.
Published: (2025)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
by: Wu, Changti, et al.
Published: (2025)
by: Wu, Changti, et al.
Published: (2025)
Personal Visual Memory from Explicit and Implicit Evidence
by: Nguyen, Viet, et al.
Published: (2026)
by: Nguyen, Viet, et al.
Published: (2026)
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
by: Liu, Tianhui, et al.
Published: (2026)
by: Liu, Tianhui, et al.
Published: (2026)
Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model
by: Huang, Kuan-Chih, et al.
Published: (2024)
by: Huang, Kuan-Chih, et al.
Published: (2024)
CIVET: Systematic Evaluation of Understanding in VLMs
by: Rizzoli, Massimo, et al.
Published: (2025)
by: Rizzoli, Massimo, et al.
Published: (2025)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
by: Jia, Mengdi, et al.
Published: (2025)
by: Jia, Mengdi, et al.
Published: (2025)
Mechanistic Diagnostics of Spatial Lexical Bias in Multimodal Large Language Model Spatial Reasoning
by: Ma, Chuang, et al.
Published: (2026)
by: Ma, Chuang, et al.
Published: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
by: Shi, Chufan, et al.
Published: (2026)
by: Shi, Chufan, et al.
Published: (2026)
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation
by: Uehara, Kohei, et al.
Published: (2024)
by: Uehara, Kohei, et al.
Published: (2024)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
by: Zhan, Xiaoyu, et al.
Published: (2025)
by: Zhan, Xiaoyu, et al.
Published: (2025)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
by: Batra, Hunar, et al.
Published: (2025)
by: Batra, Hunar, et al.
Published: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
Similar Items
-
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
by: He, Jixuan, et al.
Published: (2025) -
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
by: You, Junqi, et al.
Published: (2025) -
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025) -
Gaga: Group Any Gaussians via 3D-aware Memory Bank
by: Lyu, Weijie, et al.
Published: (2024) -
Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint
by: Zhou, Junwei, et al.
Published: (2024)