SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sohn, Tin Stribor, Dillitzer, Maximilian, Corso, Jason J., Sax, Eric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025)
Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
von: Liu, Xinhao, et al.
Veröffentlicht: (2025)
von: Liu, Xinhao, et al.
Veröffentlicht: (2025)
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)
OpenSGA: Efficient 3D Scene Graph Alignment in the Open World
von: Chen, Gang, et al.
Veröffentlicht: (2026)
von: Chen, Gang, et al.
Veröffentlicht: (2026)
Ego to World: Collaborative Spatial Reasoning in Embodied Systems via Reinforcement Learning
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
von: Zhou, Heng, et al.
Veröffentlicht: (2026)
Nav-R1: Reasoning and Navigation in Embodied Scenes
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
von: Liu, Qingxiang, et al.
Veröffentlicht: (2025)
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
von: Peng, Baorui, et al.
Veröffentlicht: (2026)
GigaWorld-0: World Models as Data Engine to Empower Embodied AI
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
von: GigaWorld Team, et al.
Veröffentlicht: (2025)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2025)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2025)
EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
von: Wang, Xinjie, et al.
Veröffentlicht: (2025)
RoboScape: Physics-informed Embodied World Model
von: Shang, Yu, et al.
Veröffentlicht: (2025)
von: Shang, Yu, et al.
Veröffentlicht: (2025)
Learning 3D Persistent Embodied World Models
von: Zhou, Siyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Siyuan, et al.
Veröffentlicht: (2025)
WoMAP: World Models For Embodied Open-Vocabulary Object Localization
von: Yin, Tenny, et al.
Veröffentlicht: (2025)
von: Yin, Tenny, et al.
Veröffentlicht: (2025)
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
von: Shang, Yu, et al.
Veröffentlicht: (2026)
von: Shang, Yu, et al.
Veröffentlicht: (2026)
WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
von: Shang, Yu, et al.
Veröffentlicht: (2026)
von: Shang, Yu, et al.
Veröffentlicht: (2026)
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
von: Yang, Yuncong, et al.
Veröffentlicht: (2024)
TesserAct: Learning 4D Embodied World Models
von: Zhen, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2025)
CLIPping the Limits: Finding the Sweet Spot for Relevant Images in Automated Driving Systems Perception Testing
von: Rigoll, Philipp, et al.
Veröffentlicht: (2024)
von: Rigoll, Philipp, et al.
Veröffentlicht: (2024)
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
von: Zhang, Qifan, et al.
Veröffentlicht: (2026)
von: Zhang, Qifan, et al.
Veröffentlicht: (2026)
WoW: Towards a World omniscient World model Through Embodied Interaction
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2025)
Open-World Panoptic Segmentation
von: Sodano, Matteo, et al.
Veröffentlicht: (2024)
von: Sodano, Matteo, et al.
Veröffentlicht: (2024)
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
von: Kong, Lingdong, et al.
Veröffentlicht: (2024)
SKIP: Sparse Keyframe Interpolation Paradigm for Efficient Embodied World Models
von: He, Ziheng, et al.
Veröffentlicht: (2026)
von: He, Ziheng, et al.
Veröffentlicht: (2026)
CounterScene: Counterfactual Causal Reasoning in Generative World Models for Safety-Critical Closed-Loop Evaluation
von: Jing, Bowen, et al.
Veröffentlicht: (2026)
von: Jing, Bowen, et al.
Veröffentlicht: (2026)
EVA: An Embodied World Model for Future Video Anticipation
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
von: Chi, Xiaowei, et al.
Veröffentlicht: (2024)
Articulated 3D Scene Graphs for Open-World Mobile Manipulation
von: Büchner, Martin, et al.
Veröffentlicht: (2026)
von: Büchner, Martin, et al.
Veröffentlicht: (2026)
World Action Models: The Next Frontier in Embodied AI
von: Wang, Siyin, et al.
Veröffentlicht: (2026)
von: Wang, Siyin, et al.
Veröffentlicht: (2026)
WorldVLM: Combining World Model Forecasting and Vision-Language Reasoning
von: Englmeier, Stefan, et al.
Veröffentlicht: (2026)
von: Englmeier, Stefan, et al.
Veröffentlicht: (2026)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
von: Xu, Wenjiang, et al.
Veröffentlicht: (2025)
RynnEC: Bringing MLLMs into Embodied World
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
Rethinking Video Generation Model for the Embodied World
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
von: Deng, Yufan, et al.
Veröffentlicht: (2026)
KeyWorld: Key Frame Reasoning Enables Effective and Efficient World Models
von: Li, Sibo, et al.
Veröffentlicht: (2025)
von: Li, Sibo, et al.
Veröffentlicht: (2025)
Kinema4D: Kinematic 4D World Modeling for Spatiotemporal Embodied Simulation
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
von: Xu, Mutian, et al.
Veröffentlicht: (2026)
SABER: A Scalable Action-Based Embodied Dataset for Real-World VLA Adaptation
von: Menga, Narsimha, et al.
Veröffentlicht: (2026)
von: Menga, Narsimha, et al.
Veröffentlicht: (2026)
Understanding Spatio-Temporal Relations in Human-Object Interaction using Pyramid Graph Convolutional Network
von: Xing, Hao, et al.
Veröffentlicht: (2024)
von: Xing, Hao, et al.
Veröffentlicht: (2024)
Embodied Spatial Intelligence: from Implicit Scene Modeling to Spatial Reasoning
von: Fang, Jiading
Veröffentlicht: (2025)
von: Fang, Jiading
Veröffentlicht: (2025)
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
von: Dang, Ronghao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025) -
Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025) -
A Framework for a Capability-driven Evaluation of Scenario Understanding for Multimodal Large Language Models in Autonomous Driving
von: Sohn, Tin Stribor, et al.
Veröffentlicht: (2025) -
Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
von: Liu, Xinhao, et al.
Veröffentlicht: (2025) -
ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
von: Zhou, Zhongyi, et al.
Veröffentlicht: (2025)