Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cao, Meng, Lin, Haokun, Li, Haoyuan, Tang, Haoran, Xu, Rongtao, An, Dong, Liu, Xue, Reid, Ian, Liang, Xiaodan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
von: Li, Haoyuan, et al.
Veröffentlicht: (2026)
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026)
von: Cao, Meng, et al.
Veröffentlicht: (2026)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
von: Tang, Haoran, et al.
Veröffentlicht: (2025)
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
von: Ma, Liang, et al.
Veröffentlicht: (2025)
von: Ma, Liang, et al.
Veröffentlicht: (2025)
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
von: Wang, Yongxin, et al.
Veröffentlicht: (2024)
World2Act: Latent Action Post-Training from World Model Dynamics
von: Vuong, An Dinh, et al.
Veröffentlicht: (2026)
von: Vuong, An Dinh, et al.
Veröffentlicht: (2026)
ImagineNav: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
von: Zhao, Xinxin, et al.
Veröffentlicht: (2024)
von: Zhao, Xinxin, et al.
Veröffentlicht: (2024)
AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis
von: Tang, Tao, et al.
Veröffentlicht: (2024)
von: Tang, Tao, et al.
Veröffentlicht: (2024)
BEVPose: Unveiling Scene Semantics through Pose-Guided Multi-Modal BEV Alignment
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
von: Hosseinzadeh, Mehdi, et al.
Veröffentlicht: (2024)
Constraint-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
von: Chen, Kehan, et al.
Veröffentlicht: (2024)
RoBridge: A Hierarchical Architecture Bridging Cognition and Execution for General Robotic Manipulation
von: Zhang, Kaidong, et al.
Veröffentlicht: (2025)
von: Zhang, Kaidong, et al.
Veröffentlicht: (2025)
Complementary Information Guided Occupancy Prediction via Multi-Level Representation Fusion
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
GS-CLIP: Gaussian Splatting for Contrastive Language-Image-3D Pretraining from Real-World Data
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
PhysGame: Uncovering Physical Commonsense Violations in Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
GLaD: Geometric Latent Distillation for Vision-Language-Action Models
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
von: Guo, Minghao, et al.
Veröffentlicht: (2025)
Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation
von: Hu, Yue, et al.
Veröffentlicht: (2025)
von: Hu, Yue, et al.
Veröffentlicht: (2025)
Contrastive Learning with Counterfactual Explanations for Radiology Report Generation
von: Li, Mingjie, et al.
Veröffentlicht: (2024)
von: Li, Mingjie, et al.
Veröffentlicht: (2024)
CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
von: Wang, Yongxin, et al.
Veröffentlicht: (2025)
InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models
von: Yan, Yu, et al.
Veröffentlicht: (2024)
von: Yan, Yu, et al.
Veröffentlicht: (2024)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
von: Han, Mingfei, et al.
Veröffentlicht: (2026)
von: Han, Mingfei, et al.
Veröffentlicht: (2026)
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
von: Wang, Teng, et al.
Veröffentlicht: (2025)
von: Wang, Teng, et al.
Veröffentlicht: (2025)
S-INF: Towards Realistic Indoor Scene Synthesis via Scene Implicit Neural Field
von: Liang, Zixi, et al.
Veröffentlicht: (2024)
von: Liang, Zixi, et al.
Veröffentlicht: (2024)
SpatialCrafter: Unleashing the Imagination of Video Diffusion Models for Scene Reconstruction from Limited Observations
von: Zhang, Songchun, et al.
Veröffentlicht: (2025)
von: Zhang, Songchun, et al.
Veröffentlicht: (2025)
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
von: Liang, Xiwen, et al.
Veröffentlicht: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
MUSE: Mamba is Efficient Multi-scale Learner for Text-video Retrieval
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
von: Tang, Haoran, et al.
Veröffentlicht: (2024)
3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
von: Xu, Rongtao, et al.
Veröffentlicht: (2025)
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
von: Zhang, Wanyue, et al.
Veröffentlicht: (2026)
Geometry-Aware Implicit Memory for Video World Models
von: Wei, Zhengxuan, et al.
Veröffentlicht: (2026)
von: Wei, Zhengxuan, et al.
Veröffentlicht: (2026)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
ManipArena: Comprehensive Real-world Evaluation of Reasoning-Oriented Generalist Robot Manipulation
von: Sun, Yu, et al.
Veröffentlicht: (2026)
von: Sun, Yu, et al.
Veröffentlicht: (2026)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
von: Wallingford, Matthew, et al.
Veröffentlicht: (2024)
von: Wallingford, Matthew, et al.
Veröffentlicht: (2024)
MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning
von: Li, Yian, et al.
Veröffentlicht: (2026)
von: Li, Yian, et al.
Veröffentlicht: (2026)
ForesightNav: Learning Scene Imagination for Efficient Exploration
von: Shah, Hardik, et al.
Veröffentlicht: (2025)
von: Shah, Hardik, et al.
Veröffentlicht: (2025)
Seeing the World through Your Eyes
von: Alzayer, Hadi, et al.
Veröffentlicht: (2023)
von: Alzayer, Hadi, et al.
Veröffentlicht: (2023)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
von: Xu, Yifu, et al.
Veröffentlicht: (2026)
von: Xu, Yifu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
von: Cao, Meng, et al.
Veröffentlicht: (2025) -
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning
von: Li, Haoyuan, et al.
Veröffentlicht: (2026) -
Order from Chaos: Physical World Understanding from Glitchy Gameplay Videos
von: Cao, Meng, et al.
Veröffentlicht: (2026) -
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
von: Cao, Meng, et al.
Veröffentlicht: (2025) -
Video Spatial Reasoning with Object-Centric 3D Rollout
von: Tang, Haoran, et al.
Veröffentlicht: (2025)