Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yibin, Xu, Wang, Zhang, Wanyue, Zhi, Helu, Huang, Jingjing, Xu, Yangbin, Sun, Yangang, Zhu, Conghui, Zhao, Tiejun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025)
by: Zhang, Wanyue, et al.
Published: (2025)
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
by: Helu, Zhi, et al.
Published: (2025)
by: Helu, Zhi, et al.
Published: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026)
by: Zhang, Wanyue, et al.
Published: (2026)
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps
by: Gao, Xiangjun, et al.
Published: (2026)
by: Gao, Xiangjun, et al.
Published: (2026)
Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning
by: Huang, Buzhen, et al.
Published: (2025)
by: Huang, Buzhen, et al.
Published: (2025)
Enhancing Object Coherence in Layout-to-Image Synthesis
by: Wang, Yibin, et al.
Published: (2023)
by: Wang, Yibin, et al.
Published: (2023)
DocCogito: Aligning Layout Cognition and Step-Level Grounded Reasoning for Document Understanding
by: Wu, Yuchuan, et al.
Published: (2026)
by: Wu, Yuchuan, et al.
Published: (2026)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator
by: Lin, Jieru, et al.
Published: (2024)
by: Lin, Jieru, et al.
Published: (2024)
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Generative Recall, Dense Reranking: Learning Multi-View Semantic IDs for Efficient Text-to-Video Retrieval
by: Zhao, Zecheng, et al.
Published: (2026)
by: Zhao, Zecheng, et al.
Published: (2026)
DocLayout-YOLO: Enhancing Document Layout Analysis through Diverse Synthetic Data and Global-to-Local Adaptive Perception
by: Zhao, Zhiyuan, et al.
Published: (2024)
by: Zhao, Zhiyuan, et al.
Published: (2024)
OpenGround: Active Cognition-based Reasoning for Open-World 3D Visual Grounding
by: Huang, Wenyuan, et al.
Published: (2025)
by: Huang, Wenyuan, et al.
Published: (2025)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
by: Zhu, Wenxin, et al.
Published: (2025)
by: Zhu, Wenxin, et al.
Published: (2025)
VIGIL: Part-Grounded Structured Reasoning for Generalizable Deepfake Detection
by: Li, Xinghan, et al.
Published: (2026)
by: Li, Xinghan, et al.
Published: (2026)
Closely Interactive Human Reconstruction with Proxemics and Physics-Guided Adaption
by: Huang, Buzhen, et al.
Published: (2024)
by: Huang, Buzhen, et al.
Published: (2024)
ReLayout: Versatile and Structure-Preserving Design Layout Editing via Relation-Aware Design Reconstruction
by: Lin, Jiawei, et al.
Published: (2026)
by: Lin, Jiawei, et al.
Published: (2026)
Spatial Diffusion for Cell Layout Generation
by: Li, Chen, et al.
Published: (2024)
by: Li, Chen, et al.
Published: (2024)
OmniDocLayout: Towards Diverse Document Layout Generation via Coarse-to-Fine LLM Learning
by: Kang, Hengrui, et al.
Published: (2025)
by: Kang, Hengrui, et al.
Published: (2025)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
Imaginarium: Vision-guided High-Quality 3D Scene Layout Generation
by: Zhu, Xiaoming, et al.
Published: (2025)
by: Zhu, Xiaoming, et al.
Published: (2025)
Two-Pass Zero-Shot Temporal-Spatial Grounding of Rare Traffic Events in Surveillance Video
by: Huang, Jiantang
Published: (2026)
by: Huang, Jiantang
Published: (2026)
MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
by: Dai, Shiqi, et al.
Published: (2025)
by: Dai, Shiqi, et al.
Published: (2025)
Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial Reasoning
by: Ran, Xingjian, et al.
Published: (2025)
by: Ran, Xingjian, et al.
Published: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
by: Huang, Yanxiang, et al.
Published: (2026)
by: Huang, Yanxiang, et al.
Published: (2026)
G$^2$VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
by: Hu, Wenbo, et al.
Published: (2025)
by: Hu, Wenbo, et al.
Published: (2025)
StructLayoutFormer:Conditional Structured Layout Generation via Structure Serialization and Disentanglement
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
Adapting Human Mesh Recovery with Vision-Language Feedback
by: Xu, Chongyang, et al.
Published: (2025)
by: Xu, Chongyang, et al.
Published: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
Rethinking High-speed Image Reconstruction Framework with Spike Camera
by: Chen, Kang, et al.
Published: (2025)
by: Chen, Kang, et al.
Published: (2025)
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
by: Guo, Chaohong, et al.
Published: (2026)
by: Guo, Chaohong, et al.
Published: (2026)
TALL: Thumbnail Layout for Deepfake Video Detection
by: Xu, Yuting, et al.
Published: (2023)
by: Xu, Yuting, et al.
Published: (2023)
Training-free Composite Scene Generation for Layout-to-Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2024)
by: Liu, Jiaqi, et al.
Published: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
by: Zhang, Mingfang, et al.
Published: (2026)
by: Zhang, Mingfang, et al.
Published: (2026)
Spatial-Conditioned Reasoning in Long-Egocentric Videos
by: Tribble, James, et al.
Published: (2026)
by: Tribble, James, et al.
Published: (2026)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognition
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
Similar Items
-
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
by: Zhang, Wanyue, et al.
Published: (2025) -
Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
by: Helu, Zhi, et al.
Published: (2025) -
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
by: Zhang, Wanyue, et al.
Published: (2026) -
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps
by: Gao, Xiangjun, et al.
Published: (2026) -
Reconstructing Close Human Interaction with Appearance and Proxemics Reasoning
by: Huang, Buzhen, et al.
Published: (2025)