Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Pingyue, Huang, Zihan, Wang, Yue, Zhang, Jieyu, Xue, Letian, Wang, Zihan, Wang, Qineng, Chandrasegaran, Keshigeyan, Zhang, Ruohan, Choi, Yejin, Krishna, Ranjay, Wu, Jiajun, Fei-Fei, Li, Li, Manling |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MindCube: Spatial Mental Modeling from Limited Views
por: Wang, Qineng, et al.
Publicado: (2025)
por: Wang, Qineng, et al.
Publicado: (2025)
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
por: Wang, Kangrui, et al.
Publicado: (2025)
por: Wang, Kangrui, et al.
Publicado: (2025)
RAGEN-2: Reasoning Collapse in Agentic RL
por: Wang, Zihan, et al.
Publicado: (2026)
por: Wang, Zihan, et al.
Publicado: (2026)
T*: Re-thinking Temporal Search for Long-Form Video Understanding
por: Ye, Jinhui, et al.
Publicado: (2025)
por: Ye, Jinhui, et al.
Publicado: (2025)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
Planning with the Views via Scene Self-Exploration
por: Wang, Kangrui, et al.
Publicado: (2026)
por: Wang, Kangrui, et al.
Publicado: (2026)
HourVideo: 1-Hour Video-Language Understanding
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2024)
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2024)
ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction
por: Wang, Qineng, et al.
Publicado: (2025)
por: Wang, Qineng, et al.
Publicado: (2025)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
por: Hong, Yining, et al.
Publicado: (2026)
por: Hong, Yining, et al.
Publicado: (2026)
Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs
por: Hong, Yining, et al.
Publicado: (2026)
por: Hong, Yining, et al.
Publicado: (2026)
GPIC: A Giant Permissive Image Corpus for Visual Generation
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2026)
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2026)
TRANSIC: Sim-to-Real Policy Transfer by Learning from Online Correction
por: Jiang, Yunfan, et al.
Publicado: (2024)
por: Jiang, Yunfan, et al.
Publicado: (2024)
Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making
por: Li, Manling, et al.
Publicado: (2024)
por: Li, Manling, et al.
Publicado: (2024)
Model Inversion Robustness: Can Transfer Learning Help?
por: Ho, Sy-Tuyen, et al.
Publicado: (2024)
por: Ho, Sy-Tuyen, et al.
Publicado: (2024)
Iterated Learning Improves Compositionality in Large Vision-Language Models
por: Zheng, Chenhao, et al.
Publicado: (2024)
por: Zheng, Chenhao, et al.
Publicado: (2024)
Exploring Diffusion Transformer Designs via Grafting
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2025)
por: Chandrasegaran, Keshigeyan, et al.
Publicado: (2025)
Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
por: Dharmarajan, Karthik, et al.
Publicado: (2025)
por: Dharmarajan, Karthik, et al.
Publicado: (2025)
I Can Tell What I am Doing: Toward Real-World Natural Language Grounding of Robot Experiences
por: Wang, Zihan, et al.
Publicado: (2024)
por: Wang, Zihan, et al.
Publicado: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
por: Li, Linjie, et al.
Publicado: (2025)
por: Li, Linjie, et al.
Publicado: (2025)
ReKep: Spatio-Temporal Reasoning of Relational Keypoint Constraints for Robotic Manipulation
por: Huang, Wenlong, et al.
Publicado: (2024)
por: Huang, Wenlong, et al.
Publicado: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
por: Gao, Ziqi, et al.
Publicado: (2024)
por: Gao, Ziqi, et al.
Publicado: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
por: Ma, Zixian, et al.
Publicado: (2024)
por: Ma, Zixian, et al.
Publicado: (2024)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
por: Yang, Yinuo, et al.
Publicado: (2026)
por: Yang, Yinuo, et al.
Publicado: (2026)
Offline Training of Language Model Agents with Functions as Learnable Weights
por: Zhang, Shaokun, et al.
Publicado: (2024)
por: Zhang, Shaokun, et al.
Publicado: (2024)
Can Large Language Models Reinvent Foundational Algorithms?
por: Zhao, Jian, et al.
Publicado: (2026)
por: Zhao, Jian, et al.
Publicado: (2026)
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
por: Zhao, Hongjue, et al.
Publicado: (2026)
por: Zhao, Hongjue, et al.
Publicado: (2026)
A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
por: Liu, Licheng, et al.
Publicado: (2025)
por: Liu, Licheng, et al.
Publicado: (2025)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
por: Tang, Yihe, et al.
Publicado: (2025)
por: Tang, Yihe, et al.
Publicado: (2025)
Non-unique decompositions of mixed states and deterministic energy transfers
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
VisionUnite: A Vision-Language Foundation Model for Ophthalmology Enhanced with Clinical Knowledge
por: Li, Zihan, et al.
Publicado: (2024)
por: Li, Zihan, et al.
Publicado: (2024)
AirShot: Efficient Few-Shot Detection for Autonomous Exploration
por: Wang, Zihan, et al.
Publicado: (2024)
por: Wang, Zihan, et al.
Publicado: (2024)
Automated Creation of Digital Cousins for Robust Policy Learning
por: Dai, Tianyuan, et al.
Publicado: (2024)
por: Dai, Tianyuan, et al.
Publicado: (2024)
Chain-of-Experts: Unlocking the Communication Power of Mixture-of-Experts Models
por: Wang, Zihan, et al.
Publicado: (2025)
por: Wang, Zihan, et al.
Publicado: (2025)
Order Matters: Rethinking Prompt Construction in In-Context Learning
por: Li, Warren, et al.
Publicado: (2025)
por: Li, Warren, et al.
Publicado: (2025)
DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation
por: Wang, Chen, et al.
Publicado: (2024)
por: Wang, Chen, et al.
Publicado: (2024)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
por: Abdollahzadeh, Milad, et al.
Publicado: (2023)
por: Abdollahzadeh, Milad, et al.
Publicado: (2023)
IMPASTO: Integrating Model-Based Planning with Learned Dynamics Models for Robotic Oil Painting Reproduction
por: Wang, Yingke, et al.
Publicado: (2026)
por: Wang, Yingke, et al.
Publicado: (2026)
EMCompress: Video-LLMs with Endomorphic Multimodal Compression
por: Fan, Zheyu, et al.
Publicado: (2025)
por: Fan, Zheyu, et al.
Publicado: (2025)
Weak-to-Strong Generalization with Failure Trajectories: A Tree-based Approach to Elicit Optimal Policy in Strong Models
por: Ye, Ruimeng, et al.
Publicado: (2025)
por: Ye, Ruimeng, et al.
Publicado: (2025)
Artificial Entanglement in the Fine-Tuning of Large Language Models
por: Chen, Min, et al.
Publicado: (2026)
por: Chen, Min, et al.
Publicado: (2026)
Ejemplares similares
-
MindCube: Spatial Mental Modeling from Limited Views
por: Wang, Qineng, et al.
Publicado: (2025) -
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
por: Wang, Kangrui, et al.
Publicado: (2025) -
RAGEN-2: Reasoning Collapse in Agentic RL
por: Wang, Zihan, et al.
Publicado: (2026) -
T*: Re-thinking Temporal Search for Long-Form Video Understanding
por: Ye, Jinhui, et al.
Publicado: (2025) -
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
por: Wang, Zihan, et al.
Publicado: (2025)