SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Yuecheng, Chi, Dafeng, Wu, Shiguang, Zhang, Zhanguang, Hu, Yaochen, Zhang, Lingfeng, Zhang, Yingxue, Wu, Shuang, Cao, Tongtong, Huang, Guowei, Huang, Helong, Tian, Guangjian, Qiu, Weichao, Quan, Xingyue, Hao, Jianye, Zhuang, Yuzheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025)
ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2024)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2024)
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
von: Cai, Xinyan, et al.
Veröffentlicht: (2025)
von: Cai, Xinyan, et al.
Veröffentlicht: (2025)
Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Following
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions
von: Huang, Helong, et al.
Veröffentlicht: (2025)
von: Huang, Helong, et al.
Veröffentlicht: (2025)
Whole-Body Inverse Kinematics with Graph Diffusion
von: Huang, Helong, et al.
Veröffentlicht: (2026)
von: Huang, Helong, et al.
Veröffentlicht: (2026)
Self-CriTeach: LLM Self-Teaching and Self-Critiquing for Improving Robotic Planning via Automated Domain Generation
von: Huang, Jinbang, et al.
Veröffentlicht: (2025)
von: Huang, Jinbang, et al.
Veröffentlicht: (2025)
H-WM: Robotic Task and Motion Planning Guided by Hierarchical World Model
von: Huang, Jinbang, et al.
Veröffentlicht: (2026)
von: Huang, Jinbang, et al.
Veröffentlicht: (2026)
One Demo Is All It Takes: Planning Domain Derivation with LLMs from A Single Demonstration
von: Huang, Jinbang, et al.
Veröffentlicht: (2025)
von: Huang, Jinbang, et al.
Veröffentlicht: (2025)
ForceFlow: Learning to Feel and Act via Contact-Driven Flow Matching
von: Zhang, Shuoheng, et al.
Veröffentlicht: (2026)
von: Zhang, Shuoheng, et al.
Veröffentlicht: (2026)
Do World Action Models Generalize Better than VLAs? A Robustness Study
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2026)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
von: Shao, Xinyu, et al.
Veröffentlicht: (2025)
von: Shao, Xinyu, et al.
Veröffentlicht: (2025)
Preference and Concurrence Aware Bayesian Graph Neural Networks for Recommender Systems
von: Gu, Hongjian, et al.
Veröffentlicht: (2023)
von: Gu, Hongjian, et al.
Veröffentlicht: (2023)
Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
von: Ni, Fei, et al.
Veröffentlicht: (2025)
von: Ni, Fei, et al.
Veröffentlicht: (2025)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
EmbodiedVSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
von: Du, Mengfei, et al.
Veröffentlicht: (2024)
HardCore Generation: Generating Hard UNSAT Problems for Data Augmentation
von: Cotnareanu, Joseph, et al.
Veröffentlicht: (2024)
von: Cotnareanu, Joseph, et al.
Veröffentlicht: (2024)
SpatialPoint: Spatial-aware Point Prediction for Embodied Localization
von: Zhu, Qiming, et al.
Veröffentlicht: (2026)
von: Zhu, Qiming, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
von: Gao, Jianzhe, et al.
Veröffentlicht: (2026)
The Graph's Apprentice: Teaching an LLM Low Level Knowledge for Circuit Quality Estimation
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
von: Moravej, Reza, et al.
Veröffentlicht: (2024)
E2ESlack: An End-to-End Graph-Based Framework for Pre-Routing Slack Prediction
von: Bodhe, Saurabh, et al.
Veröffentlicht: (2025)
von: Bodhe, Saurabh, et al.
Veröffentlicht: (2025)
MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
von: Wei, Ziming, et al.
Veröffentlicht: (2025)
Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Framework
von: Huang, Kerui, et al.
Veröffentlicht: (2025)
von: Huang, Kerui, et al.
Veröffentlicht: (2025)
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models
von: Zhang, Jiyao, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyao, et al.
Veröffentlicht: (2026)
GraSS: Combining Graph Neural Networks with Expert Knowledge for SAT Solver Selection
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhanguang, et al.
Veröffentlicht: (2024)
SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
von: Wu, Haoning, et al.
Veröffentlicht: (2025)
von: Wu, Haoning, et al.
Veröffentlicht: (2025)
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
von: Lin, Xiangye, et al.
Veröffentlicht: (2026)
von: Lin, Xiangye, et al.
Veröffentlicht: (2026)
Maffei's action and symplectic Springer action for quiver varieties
von: Wu, Yaochen
Veröffentlicht: (2024)
von: Wu, Yaochen
Veröffentlicht: (2024)
Enhancing Logical Reasoning in Large Language Models through Graph-based Synthetic Data
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
von: Alomrani, Mohammad Ali, et al.
Veröffentlicht: (2025)
von: Alomrani, Mohammad Ali, et al.
Veröffentlicht: (2025)
Grid Spatial Understanding: A Dataset for Textual Spatial Reasoning over Grids, Embodied Settings, and Coordinate Structures
von: Sidhu, Risham, et al.
Veröffentlicht: (2026)
von: Sidhu, Risham, et al.
Veröffentlicht: (2026)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
STEMS: Spatial-Temporal Enhanced Safe Multi-Agent Coordination for Building Energy Management
von: Zhang, Huiliang, et al.
Veröffentlicht: (2025)
von: Zhang, Huiliang, et al.
Veröffentlicht: (2025)
Efficient Reasoning for LLMs through Speculative Chain-of-Thought
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
von: Wang, Jikai, et al.
Veröffentlicht: (2025)
CoFlow: Coordinated Few-Step Flow for Offline Multi-Agent Decision Making
von: Zou, Guowei, et al.
Veröffentlicht: (2026)
von: Zou, Guowei, et al.
Veröffentlicht: (2026)
Extracting and Following Paths for Robust Relational Reasoning with Large Language Models
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
von: Li, Jiangyang, et al.
Veröffentlicht: (2026)
von: Li, Jiangyang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
von: Liu, Yuecheng, et al.
Veröffentlicht: (2025) -
ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2024) -
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025) -
EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
von: Cai, Xinyan, et al.
Veröffentlicht: (2025) -
Astra: Efficient Transformer Architecture and Contrastive Dynamics Learning for Embodied Instruction Following
von: Ma, Yueen, et al.
Veröffentlicht: (2024)