Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Heng, Kang, Li, Qin, Yiran, Song, Xiufeng, Yu, Ao, Zhang, Zilu, Song, Haoming, Xu, Kaixin, Fan, Yuchen, Zhou, Dongzhan, Liu, Xiaohong, Zhang, Ruimao, Torr, Philip, Bai, Lei, Yin, Zhenfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
von: Kang, Li, et al.
Veröffentlicht: (2025)
von: Kang, Li, et al.
Veröffentlicht: (2025)
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
von: Kang, Li, et al.
Veröffentlicht: (2026)
von: Kang, Li, et al.
Veröffentlicht: (2026)
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
von: Qin, Yiran, et al.
Veröffentlicht: (2026)
von: Qin, Yiran, et al.
Veröffentlicht: (2026)
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
von: Qin, Yiran, et al.
Veröffentlicht: (2023)
Select-then-Solve: Paradigm Routing as Inference-Time Optimization for LLM Agents
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
von: Zhou, Enshen, et al.
Veröffentlicht: (2024)
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
von: Fan, Kaixuan, et al.
Veröffentlicht: (2025)
von: Fan, Kaixuan, et al.
Veröffentlicht: (2025)
NavigateDiff: Visual Predictors are Zero-Shot Navigation Assistants
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
von: Qin, Yiran, et al.
Veröffentlicht: (2025)
WorldSimBench: Towards Video Generation Models as World Simulators
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
von: Qin, Yiran, et al.
Veröffentlicht: (2024)
LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge
von: Zhou, Heng, et al.
Veröffentlicht: (2025)
von: Zhou, Heng, et al.
Veröffentlicht: (2025)
ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
von: Zhou, Heng, et al.
Veröffentlicht: (2025)
von: Zhou, Heng, et al.
Veröffentlicht: (2025)
GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion Policies
von: Wang, Ziye, et al.
Veröffentlicht: (2025)
von: Wang, Ziye, et al.
Veröffentlicht: (2025)
High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian Representation
von: Wang, Ziye, et al.
Veröffentlicht: (2025)
von: Wang, Ziye, et al.
Veröffentlicht: (2025)
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
von: Tan, Zelin, et al.
Veröffentlicht: (2025)
Sentinel: Embodied Cooperative Spatial Reasoning and Planning
von: Lin, Xiangye, et al.
Veröffentlicht: (2026)
von: Lin, Xiangye, et al.
Veröffentlicht: (2026)
Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
von: Zhao, Baining, et al.
Veröffentlicht: (2025)
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
von: Gholami, Mohsen, et al.
Veröffentlicht: (2025)
World-Ego Modeling for Long-Horizon Evolution in Hybrid Embodied Tasks
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
von: Lin, Zuyao, et al.
Veröffentlicht: (2026)
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyang, et al.
Veröffentlicht: (2026)
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2026)
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2026)
Toward Accurate Camera-based 3D Object Detection via Cascade Depth Estimation and Calibration
von: Wang, Chaoqun, et al.
Veröffentlicht: (2024)
von: Wang, Chaoqun, et al.
Veröffentlicht: (2024)
Rethinking Vision-Language Model in Face Forensics: Multi-Modal Interpretable Forged Face Detector
von: Guo, Xiao, et al.
Veröffentlicht: (2025)
von: Guo, Xiao, et al.
Veröffentlicht: (2025)
MosaicThinker: On-Device Visual Spatial Reasoning for Embodied AI via Iterative Construction of Space Representation
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
von: Wang, Haoming, et al.
Veröffentlicht: (2026)
Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?
von: Yu, Junchi, et al.
Veröffentlicht: (2025)
von: Yu, Junchi, et al.
Veröffentlicht: (2025)
Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
von: Liu, Xiaohong, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohong, et al.
Veröffentlicht: (2025)
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
von: Hao, Jinkun, et al.
Veröffentlicht: (2026)
LabUtopia: High-Fidelity Simulation and Hierarchical Benchmark for Scientific Embodied Agents
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion
von: Ma, Jiahua, et al.
Veröffentlicht: (2025)
von: Ma, Jiahua, et al.
Veröffentlicht: (2025)
TouchGuide: Inference-Time Steering of Visuomotor Policies via Touch Guidance
von: Zhang, Zhemeng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhemeng, et al.
Veröffentlicht: (2026)
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materials
von: Lv, Taoyuze, et al.
Veröffentlicht: (2025)
von: Lv, Taoyuze, et al.
Veröffentlicht: (2025)
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
On Learning Multi-Modal Forgery Representation for Diffusion Generated Video Detection
von: Song, Xiufeng, et al.
Veröffentlicht: (2024)
von: Song, Xiufeng, et al.
Veröffentlicht: (2024)
Reinforcing VLAs in Task-Agnostic World Models
von: Wang, Yucen, et al.
Veröffentlicht: (2026)
von: Wang, Yucen, et al.
Veröffentlicht: (2026)
ReCrit: Transition-Aware Reinforcement Learning for Scientific Critic Reasoning
von: Xu, Wanghan, et al.
Veröffentlicht: (2026)
von: Xu, Wanghan, et al.
Veröffentlicht: (2026)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2025)
von: Xue, Xiangyuan, et al.
Veröffentlicht: (2025)
Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
von: Zhang, Lingfeng, et al.
Veröffentlicht: (2025)
Position: Intelligent Science Laboratory Requires the Integration of Cognitive and Embodied AI
von: Zhang, Sha, et al.
Veröffentlicht: (2025)
von: Zhang, Sha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
von: Qin, Yiran, et al.
Veröffentlicht: (2025) -
VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
von: Kang, Li, et al.
Veröffentlicht: (2025) -
CoEnv: Driving Embodied Multi-Agent Collaboration via Compositional Environment
von: Kang, Li, et al.
Veröffentlicht: (2026) -
ComSim: Building Scalable Real-World Robot Data Generation via Compositional Simulation
von: Qin, Yiran, et al.
Veröffentlicht: (2026) -
MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception
von: Qin, Yiran, et al.
Veröffentlicht: (2023)