Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Fangrui, Wang, Hanhui, Xie, Yiming, Gu, Jing, Ding, Tianye, Yang, Jianwei, Jiang, Huaizu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Flexible Visual Relationship Segmentation
by: Zhu, Fangrui, et al.
Published: (2024)
by: Zhu, Fangrui, et al.
Published: (2024)
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026)
by: Goswami, Prajnan, et al.
Published: (2026)
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025)
by: Ding, Tianye, et al.
Published: (2025)
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024)
by: Ding, Tianye, et al.
Published: (2024)
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023)
by: Han, Zeyu, et al.
Published: (2023)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
by: Yao, Chun-Han, et al.
Published: (2025)
by: Yao, Chun-Han, et al.
Published: (2025)
A Strong Baseline for Point Cloud Registration via Direct Superpoints Matching
by: Gupta, Aniket, et al.
Published: (2023)
by: Gupta, Aniket, et al.
Published: (2023)
SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
by: Xie, Yiming, et al.
Published: (2024)
by: Xie, Yiming, et al.
Published: (2024)
EgoReasoner: Learning Egocentric 4D Reasoning via Task-Adaptive Structured Thinking
by: Zhu, Fangrui, et al.
Published: (2026)
by: Zhu, Fangrui, et al.
Published: (2026)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
by: Meng, Zichong, et al.
Published: (2024)
by: Meng, Zichong, et al.
Published: (2024)
Absolute Coordinates Make Motion Generation Easy
by: Meng, Zichong, et al.
Published: (2025)
by: Meng, Zichong, et al.
Published: (2025)
SNAP: Towards Segmenting Anything in Any Point Cloud
by: Gupta, Aniket, et al.
Published: (2025)
by: Gupta, Aniket, et al.
Published: (2025)
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024)
by: Zhong, Lei, et al.
Published: (2024)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
by: Zhu, Rui, et al.
Published: (2026)
by: Zhu, Rui, et al.
Published: (2026)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Reasoning Portability: Guiding Continual Learning for MLLMs in the RLVR Era
by: Hong, Qiuhe, et al.
Published: (2026)
by: Hong, Qiuhe, et al.
Published: (2026)
DCVNet: Dilated Cost Volume Networks for Fast Optical Flow
by: Jiang, Huaizu, et al.
Published: (2021)
by: Jiang, Huaizu, et al.
Published: (2021)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
by: Nguyen, Hieu T., et al.
Published: (2024)
by: Nguyen, Hieu T., et al.
Published: (2024)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
by: Ouyang, Kun, et al.
Published: (2025)
by: Ouyang, Kun, et al.
Published: (2025)
GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing
by: Sun, Qigan, et al.
Published: (2026)
by: Sun, Qigan, et al.
Published: (2026)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
S$^2$-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
by: Jiang, Yankai, et al.
Published: (2026)
by: Jiang, Yankai, et al.
Published: (2026)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
by: Du, Yuetian, et al.
Published: (2026)
by: Du, Yuetian, et al.
Published: (2026)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
by: Chen, Zhenghao, et al.
Published: (2026)
by: Chen, Zhenghao, et al.
Published: (2026)
3D Spatial Understanding in MLLMs: Disambiguation and Evaluation
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
by: Guo, Longteng, et al.
Published: (2026)
by: Guo, Longteng, et al.
Published: (2026)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
by: Lu, Junxin, et al.
Published: (2026)
by: Lu, Junxin, et al.
Published: (2026)
Touch-R1: Reinforcing Touch Reasoning in MLLMs
by: Lai, Yingxin, et al.
Published: (2026)
by: Lai, Yingxin, et al.
Published: (2026)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
by: Li, Qiaoru, et al.
Published: (2026)
by: Li, Qiaoru, et al.
Published: (2026)
Spatial Preference Rewarding for MLLMs Spatial Understanding
by: Qiu, Han, et al.
Published: (2025)
by: Qiu, Han, et al.
Published: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
SpatialReward: Bridging the Perception Gap in Online RL for Image Editing via Explicit Spatial Reasoning
by: Long, Yancheng, et al.
Published: (2026)
by: Long, Yancheng, et al.
Published: (2026)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
by: Gu, Bo, et al.
Published: (2026)
by: Gu, Bo, et al.
Published: (2026)
Similar Items
-
Towards Flexible Visual Relationship Segmentation
by: Zhu, Fangrui, et al.
Published: (2024) -
UniCorrn: Unified Correspondence Transformer Across 2D and 3D
by: Goswami, Prajnan, et al.
Published: (2026) -
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
by: Ding, Tianye, et al.
Published: (2025) -
ODTFormer: Efficient Obstacle Detection and Tracking with Stereo Cameras Based on Transformer
by: Ding, Tianye, et al.
Published: (2024) -
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023)