SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Gu, Bo, Zhang, Zhikang, Wei, Zizhuang, Chen, Zhenyuan, Li, Lingyun, Song, Zhuoyi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
di: Zhao, Ruosen, et al.
Pubblicazione: (2025)
di: Zhao, Ruosen, et al.
Pubblicazione: (2025)
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
di: Chen, Zhenghao, et al.
Pubblicazione: (2026)
di: Chen, Zhenghao, et al.
Pubblicazione: (2026)
Token Entropy Regularization for Multi-modal Antenna Affiliation Identification
di: Chen, Dong, et al.
Pubblicazione: (2026)
di: Chen, Dong, et al.
Pubblicazione: (2026)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
di: Ouyang, Kun, et al.
Pubblicazione: (2025)
di: Ouyang, Kun, et al.
Pubblicazione: (2025)
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
di: Huang, Yibin, et al.
Pubblicazione: (2025)
di: Huang, Yibin, et al.
Pubblicazione: (2025)
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
di: Wang, Hengyi, et al.
Pubblicazione: (2026)
di: Wang, Hengyi, et al.
Pubblicazione: (2026)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
di: Wang, Siting, et al.
Pubblicazione: (2025)
di: Wang, Siting, et al.
Pubblicazione: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
di: Plizzari, Chiara, et al.
Pubblicazione: (2024)
di: Plizzari, Chiara, et al.
Pubblicazione: (2024)
Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language Models
di: Jang, Jaeyun, et al.
Pubblicazione: (2026)
di: Jang, Jaeyun, et al.
Pubblicazione: (2026)
PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
di: Wu, Yixuan, et al.
Pubblicazione: (2025)
di: Wu, Yixuan, et al.
Pubblicazione: (2025)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
di: Alansari, Mohamad, et al.
Pubblicazione: (2026)
LiDAR-based Registration against Georeferenced Models for Globally Consistent Allocentric Maps
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
di: Quenzel, Jan, et al.
Pubblicazione: (2024)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis
di: Zhong, Yong, et al.
Pubblicazione: (2025)
di: Zhong, Yong, et al.
Pubblicazione: (2025)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
di: Fang, Bo, et al.
Pubblicazione: (2025)
di: Fang, Bo, et al.
Pubblicazione: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
di: Qiu, Han, et al.
Pubblicazione: (2025)
di: Qiu, Han, et al.
Pubblicazione: (2025)
Towards Long-Form Spatio-Temporal Video Grounding
di: Gu, Xin, et al.
Pubblicazione: (2026)
di: Gu, Xin, et al.
Pubblicazione: (2026)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
di: Zhu, Rui, et al.
Pubblicazione: (2026)
di: Zhu, Rui, et al.
Pubblicazione: (2026)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
di: Li, Yunheng, et al.
Pubblicazione: (2026)
di: Li, Yunheng, et al.
Pubblicazione: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
di: Tasnim, Nazia, et al.
Pubblicazione: (2025)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
di: Xiao, Yuxi, et al.
Pubblicazione: (2025)
di: Xiao, Yuxi, et al.
Pubblicazione: (2025)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
di: Li, Yun, et al.
Pubblicazione: (2025)
di: Li, Yun, et al.
Pubblicazione: (2025)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
di: Li, Rang, et al.
Pubblicazione: (2025)
di: Li, Rang, et al.
Pubblicazione: (2025)
Towards Unified Surgical Scene Understanding:Bridging Reasoning and Grounding via MLLMs
di: Huang, Jincai, et al.
Pubblicazione: (2026)
di: Huang, Jincai, et al.
Pubblicazione: (2026)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
di: Guan, Tongkun, et al.
Pubblicazione: (2026)
di: Guan, Tongkun, et al.
Pubblicazione: (2026)
Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary
di: Tasnim, Nazia, et al.
Pubblicazione: (2026)
di: Tasnim, Nazia, et al.
Pubblicazione: (2026)
Exploring the Design Space of Visual Context Representation in Video MLLMs
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
Unified Representation Space for 3D Visual Grounding
di: Zheng, Yinuo, et al.
Pubblicazione: (2025)
di: Zheng, Yinuo, et al.
Pubblicazione: (2025)
Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition
di: Zhang, Xiaoying, et al.
Pubblicazione: (2025)
di: Zhang, Xiaoying, et al.
Pubblicazione: (2025)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
di: Zhao, Jiahe, et al.
Pubblicazione: (2025)
Are Image-to-Video Models Good Zero-Shot Image Editors?
di: Zhang, Zechuan, et al.
Pubblicazione: (2025)
di: Zhang, Zechuan, et al.
Pubblicazione: (2025)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
di: Yao, Jiali, et al.
Pubblicazione: (2025)
di: Yao, Jiali, et al.
Pubblicazione: (2025)
RPBG: Towards Robust Neural Point-based Graphics in the Wild
di: Zhu, Qingtian, et al.
Pubblicazione: (2024)
di: Zhu, Qingtian, et al.
Pubblicazione: (2024)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
di: Anvekar, Tejas, et al.
Pubblicazione: (2025)
di: Anvekar, Tejas, et al.
Pubblicazione: (2025)
RieMind: Geometry-Grounded Spatial Agent for Scene Understanding
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
di: Ropero, Fernando, et al.
Pubblicazione: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
di: Feng, Kaituo, et al.
Pubblicazione: (2025)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
di: Liu, Shizhan, et al.
Pubblicazione: (2025)
di: Liu, Shizhan, et al.
Pubblicazione: (2025)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
di: Zhu, Fangrui, et al.
Pubblicazione: (2025)
di: Zhu, Fangrui, et al.
Pubblicazione: (2025)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
di: Zhang, Bob, et al.
Pubblicazione: (2025)
di: Zhang, Bob, et al.
Pubblicazione: (2025)
Multi-sentence Video Grounding for Long Video Generation
di: Feng, Wei, et al.
Pubblicazione: (2024)
di: Feng, Wei, et al.
Pubblicazione: (2024)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
di: Yang, Zhuoyi, et al.
Pubblicazione: (2026)
di: Yang, Zhuoyi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
di: Zhao, Ruosen, et al.
Pubblicazione: (2025) -
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
di: Chen, Zhenghao, et al.
Pubblicazione: (2026) -
Token Entropy Regularization for Multi-modal Antenna Affiliation Identification
di: Chen, Dong, et al.
Pubblicazione: (2026) -
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
di: Ouyang, Kun, et al.
Pubblicazione: (2025) -
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
di: Huang, Yibin, et al.
Pubblicazione: (2025)