EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Zhenghao, Wang, Huiqun, Huang, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
von: Li, Jinzhao, et al.
Veröffentlicht: (2026)
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
von: Gu, Bo, et al.
Veröffentlicht: (2026)
von: Gu, Bo, et al.
Veröffentlicht: (2026)
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
von: He, Yuping, et al.
Veröffentlicht: (2025)
von: He, Yuping, et al.
Veröffentlicht: (2025)
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
von: Zhou, Nan, et al.
Veröffentlicht: (2026)
von: Zhou, Nan, et al.
Veröffentlicht: (2026)
Implicit Modeling for Transferability Estimation of Vision Foundation Models
von: Zheng, Yaoyan, et al.
Veröffentlicht: (2025)
von: Zheng, Yaoyan, et al.
Veröffentlicht: (2025)
Video-MSR: Benchmarking Multi-hop Spatial Reasoning Capabilities of MLLMs
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
von: Zhu, Rui, et al.
Veröffentlicht: (2026)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
von: Luo, Meng, et al.
Veröffentlicht: (2026)
von: Luo, Meng, et al.
Veröffentlicht: (2026)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
von: Dai, Yang, et al.
Veröffentlicht: (2026)
von: Dai, Yang, et al.
Veröffentlicht: (2026)
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
von: Zhu, Fangrui, et al.
Veröffentlicht: (2025)
Lang2Act: Fine-Grained Visual Reasoning through Self-Emergent Linguistic Toolchains
von: Xiong, Yuqi, et al.
Veröffentlicht: (2026)
von: Xiong, Yuqi, et al.
Veröffentlicht: (2026)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
Training-Free Reasoning and Reflection in MLLMs
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
von: Wei, Hongchen, et al.
Veröffentlicht: (2025)
SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
von: Wang, Peiyao, et al.
Veröffentlicht: (2025)
Emotional Theory of Mind: Bridging Fast Visual Processing with Slow Linguistic Reasoning
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
von: Etesam, Yasaman, et al.
Veröffentlicht: (2023)
RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
von: He, Jun, et al.
Veröffentlicht: (2026)
von: He, Jun, et al.
Veröffentlicht: (2026)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
von: Wang, Siting, et al.
Veröffentlicht: (2025)
von: Wang, Siting, et al.
Veröffentlicht: (2025)
Spatial Preference Rewarding for MLLMs Spatial Understanding
von: Qiu, Han, et al.
Veröffentlicht: (2025)
von: Qiu, Han, et al.
Veröffentlicht: (2025)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
von: Pei, Baoqi, et al.
Veröffentlicht: (2025)
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind
von: Plizzari, Chiara, et al.
Veröffentlicht: (2024)
von: Plizzari, Chiara, et al.
Veröffentlicht: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
From Recognition to Reasoning: Benchmarking and Enhancing MLLMs on Real-World Receipt Document Understanding
von: Wang, Yandi, et al.
Veröffentlicht: (2026)
von: Wang, Yandi, et al.
Veröffentlicht: (2026)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
von: Chen, Mingrui, et al.
Veröffentlicht: (2026)
SpatialTree: How Spatial Abilities Branch Out in MLLMs
von: Xiao, Yuxi, et al.
Veröffentlicht: (2025)
von: Xiao, Yuxi, et al.
Veröffentlicht: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
PCIE_EgoHandPose Solution for EgoExo4D Hand Pose Challenge
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
von: Li, Kailing, et al.
Veröffentlicht: (2025)
von: Li, Kailing, et al.
Veröffentlicht: (2025)
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
von: Huang, Yibin, et al.
Veröffentlicht: (2025)
von: Huang, Yibin, et al.
Veröffentlicht: (2025)
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
von: Anand, Dhruv, et al.
Veröffentlicht: (2025)
Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
von: Guo, Longteng, et al.
Veröffentlicht: (2026)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
von: Zhang, Wanyue, et al.
Veröffentlicht: (2025)
von: Zhang, Wanyue, et al.
Veröffentlicht: (2025)
Radar and Event Camera Fusion for Agile Robot Ego-Motion Estimation
von: Lyu, Yang, et al.
Veröffentlicht: (2025)
von: Lyu, Yang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
von: Li, Jinzhao, et al.
Veröffentlicht: (2026) -
SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
von: Gu, Bo, et al.
Veröffentlicht: (2026) -
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
von: He, Yuping, et al.
Veröffentlicht: (2025) -
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
von: Ouyang, Kun, et al.
Veröffentlicht: (2025) -
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)