Lifting Embodied World Models for Planning and Control
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Alex N., Darrell, Trevor, Izmailov, Pavel, Bai, Yutong, Bar, Amir |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025)
by: Bai, Yutong, et al.
Published: (2025)
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025)
by: Kutscher, Declan, et al.
Published: (2025)
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024)
by: Luo, Grace, et al.
Published: (2024)
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)
by: Bar, Amir, et al.
Published: (2023)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Analyzing The Language of Visual Tokens
by: Chan, David M., et al.
Published: (2024)
by: Chan, David M., et al.
Published: (2024)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
by: Yu, Junwei, et al.
Published: (2025)
by: Yu, Junwei, et al.
Published: (2025)
Hidden in plain sight: VLMs overlook their visual representations
by: Fu, Stephanie, et al.
Published: (2025)
by: Fu, Stephanie, et al.
Published: (2025)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
by: Mitra, Chancharik, et al.
Published: (2023)
by: Mitra, Chancharik, et al.
Published: (2023)
Visual Lexicon: Rich Image Features in Language Space
by: Wang, XuDong, et al.
Published: (2024)
by: Wang, XuDong, et al.
Published: (2024)
Finding Visual Task Vectors
by: Hojel, Alberto, et al.
Published: (2024)
by: Hojel, Alberto, et al.
Published: (2024)
3D-LFM: Lifting Foundation Model
by: Dabhi, Mosam, et al.
Published: (2023)
by: Dabhi, Mosam, et al.
Published: (2023)
An Embodied Generalist Agent in 3D World
by: Huang, Jiangyong, et al.
Published: (2023)
by: Huang, Jiangyong, et al.
Published: (2023)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
by: Qin, Yiming, et al.
Published: (2025)
by: Qin, Yiming, et al.
Published: (2025)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
by: Shang, Chuyi, et al.
Published: (2024)
by: Shang, Chuyi, et al.
Published: (2024)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
by: Lee, Heekyung, et al.
Published: (2025)
by: Lee, Heekyung, et al.
Published: (2025)
Improving Compositional Generation with Diffusion Models Using Lift Scores
by: Yu, Chenning, et al.
Published: (2025)
by: Yu, Chenning, et al.
Published: (2025)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
PAIR-Diffusion: A Comprehensive Multimodal Object-Level Image Editor
by: Goel, Vidit, et al.
Published: (2023)
by: Goel, Vidit, et al.
Published: (2023)
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024)
by: Petryk, Suzanne, et al.
Published: (2024)
Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning
by: Huang, Brandon, et al.
Published: (2024)
by: Huang, Brandon, et al.
Published: (2024)
ICAT: Incident-Case-Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models
by: Lai, Zhenglin, et al.
Published: (2026)
by: Lai, Zhenglin, et al.
Published: (2026)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
Diverse Subset Selection via Norm-Based Sampling and Orthogonality
by: Bar, Noga, et al.
Published: (2024)
by: Bar, Noga, et al.
Published: (2024)
AmaraSpatial-10K: A Spatially and Semantically Aligned 3D Dataset for Spatial Computing and Embodied AI
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
by: Salehi, Mohammad Sadegh, et al.
Published: (2026)
EmbodiedOcc: Embodied 3D Occupancy Prediction for Vision-based Online Scene Understanding
by: Wu, Yuqi, et al.
Published: (2024)
by: Wu, Yuqi, et al.
Published: (2024)
Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
by: Xu, Wenjiang, et al.
Published: (2025)
by: Xu, Wenjiang, et al.
Published: (2025)
MultiPLY: A Multisensory Object-Centric Embodied Large Language Model in 3D World
by: Hong, Yining, et al.
Published: (2024)
by: Hong, Yining, et al.
Published: (2024)
Deep Active Learning in the Open World
by: Xie, Tian, et al.
Published: (2024)
by: Xie, Tian, et al.
Published: (2024)
RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
by: Qin, Yiran, et al.
Published: (2025)
by: Qin, Yiran, et al.
Published: (2025)
World Modeling with Probabilistic Structure Integration
by: Kotar, Klemen, et al.
Published: (2025)
by: Kotar, Klemen, et al.
Published: (2025)
ExoPredicator: Learning Abstract Models of Dynamic Worlds for Robot Planning
by: Liang, Yichao, et al.
Published: (2025)
by: Liang, Yichao, et al.
Published: (2025)
TD-MPC2: Scalable, Robust World Models for Continuous Control
by: Hansen, Nicklas, et al.
Published: (2023)
by: Hansen, Nicklas, et al.
Published: (2023)
Vidar: Embodied Video Diffusion Model for Generalist Manipulation
by: Feng, Yao, et al.
Published: (2025)
by: Feng, Yao, et al.
Published: (2025)
Evaluation and Analysis of Deep Neural Transformers and Convolutional Neural Networks on Modern Remote Sensing Datasets
by: Hurt, J. Alex, et al.
Published: (2025)
by: Hurt, J. Alex, et al.
Published: (2025)
Vector Quantized Feature Fields for Fast 3D Semantic Lifting
by: Tang, George, et al.
Published: (2025)
by: Tang, George, et al.
Published: (2025)
VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
by: Liang, Yichao, et al.
Published: (2024)
by: Liang, Yichao, et al.
Published: (2024)
Similar Items
-
Navigation World Models
by: Bar, Amir, et al.
Published: (2024) -
Whole-Body Conditioned Egocentric Video Prediction
by: Bai, Yutong, et al.
Published: (2025) -
REOrdering Patches Improves Vision Models
by: Kutscher, Declan, et al.
Published: (2025) -
Vision-Language Models Create Cross-Modal Task Representations
by: Luo, Grace, et al.
Published: (2024) -
Stochastic positional embeddings improve masked image modeling
by: Bar, Amir, et al.
Published: (2023)