MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Haoming, Xue, Qiyao, Liu, Weichen, Gao, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
von: Xue, Qiyao, et al.
Veröffentlicht: (2025)
von: Xue, Qiyao, et al.
Veröffentlicht: (2025)
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
von: Xue, Qiyao, et al.
Veröffentlicht: (2024)
von: Xue, Qiyao, et al.
Veröffentlicht: (2024)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
von: Batra, Hunar, et al.
Veröffentlicht: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
von: Xiao, Zilin, et al.
Veröffentlicht: (2025)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning
von: Xu, Huilin, et al.
Veröffentlicht: (2025)
von: Xu, Huilin, et al.
Veröffentlicht: (2025)
Unified Thinker: A General Reasoning Modular Core for Image Generation
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
von: Zhou, Sashuai, et al.
Veröffentlicht: (2026)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
von: Chen, Kaitao, et al.
Veröffentlicht: (2025)
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing
von: Yang, Hanqing, et al.
Veröffentlicht: (2026)
von: Yang, Hanqing, et al.
Veröffentlicht: (2026)
SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
Do Diffusion Models Learn Semantically Meaningful and Efficient Representations?
von: Liang, Qiyao, et al.
Veröffentlicht: (2024)
von: Liang, Qiyao, et al.
Veröffentlicht: (2024)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
Learning Sparse Visual Representations via Spatial-Semantic Factorization
von: Zhao, Theodore Zhengde, et al.
Veröffentlicht: (2026)
von: Zhao, Theodore Zhengde, et al.
Veröffentlicht: (2026)
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
von: Ma, Xueqi, et al.
Veröffentlicht: (2026)
von: Ma, Xueqi, et al.
Veröffentlicht: (2026)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
DIVER: Dynamic Iterative Visual Evidence Reasoning for Multimodal Fake News Detection
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
von: Zhou, Weilin, et al.
Veröffentlicht: (2026)
StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning
von: He, Xixiang, et al.
Veröffentlicht: (2026)
von: He, Xixiang, et al.
Veröffentlicht: (2026)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
Improved Visual-Spatial Reasoning via R1-Zero-Like Training
von: Liao, Zhenyi, et al.
Veröffentlicht: (2025)
von: Liao, Zhenyi, et al.
Veröffentlicht: (2025)
Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
von: Liu, Weichen, et al.
Veröffentlicht: (2025)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
von: Zhu, Qiming, et al.
Veröffentlicht: (2026)
von: Zhu, Qiming, et al.
Veröffentlicht: (2026)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
von: Wang, Haoming, et al.
Veröffentlicht: (2025)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2024)
von: Jaiswal, Shantanu, et al.
Veröffentlicht: (2024)
TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
von: Liu, Jiahang, et al.
Veröffentlicht: (2025)
ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
von: Xue, Kaiwen, et al.
Veröffentlicht: (2026)
FloNa: Floor Plan Guided Embodied Visual Navigation
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
von: Li, Jiaxin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Reasoning Path and Latent State Analysis for Multi-view Visual Spatial Reasoning: A Cognitive Science Perspective
von: Xue, Qiyao, et al.
Veröffentlicht: (2025) -
InfiniBench: Infinite Benchmarking for Visual Spatial Reasoning with Customizable Scene Complexity
von: Wang, Haoming, et al.
Veröffentlicht: (2025) -
PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
von: Xue, Qiyao, et al.
Veröffentlicht: (2024) -
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
von: Xu, Haoran, et al.
Veröffentlicht: (2026) -
ProGait: A Multi-Purpose Video Dataset and Benchmark for Transfemoral Prosthesis Users
von: Yin, Xiangyu, et al.
Veröffentlicht: (2025)