Aether: Geometric-Aware Unified World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912503550705664 |
|---|---|
| author | Aether Team Zhu, Haoyi Wang, Yifan Zhou, Jianjun Chang, Wenzheng Zhou, Yang Li, Zizun Chen, Junyi Shen, Chunhua Pang, Jiangmiao He, Tong |
| author_facet | Aether Team Zhu, Haoyi Wang, Yifan Zhou, Jianjun Chang, Wenzheng Zhou, Yang Li, Zizun Chen, Junyi Shen, Chunhua Pang, Jiangmiao He, Tong |
| contents | The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware reasoning in world models by jointly optimizing three core capabilities: (1) 4D dynamic reconstruction, (2) action-conditioned video prediction, and (3) goal-conditioned visual planning. Through task-interleaved feature learning, Aether achieves synergistic knowledge sharing across reconstruction, prediction, and planning objectives. Building upon video generation models, our framework demonstrates zero-shot synthetic-to-real generalization despite never observing real-world data during training. Furthermore, our approach achieves zero-shot generalization in both action following and reconstruction tasks, thanks to its intrinsic geometric modeling. Notably, even without real-world data, its reconstruction performance is comparable with or even better than that of domain-specific models. Additionally, Aether employs camera trajectories as geometry-informed action spaces, enabling effective action-conditioned prediction and visual planning. We hope our work inspires the community to explore new frontiers in physically-reasonable world modeling and its applications. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_18945 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Aether: Geometric-Aware Unified World Modeling Aether Team Zhu, Haoyi Wang, Yifan Zhou, Jianjun Chang, Wenzheng Zhou, Yang Li, Zizun Chen, Junyi Shen, Chunhua Pang, Jiangmiao He, Tong Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware reasoning in world models by jointly optimizing three core capabilities: (1) 4D dynamic reconstruction, (2) action-conditioned video prediction, and (3) goal-conditioned visual planning. Through task-interleaved feature learning, Aether achieves synergistic knowledge sharing across reconstruction, prediction, and planning objectives. Building upon video generation models, our framework demonstrates zero-shot synthetic-to-real generalization despite never observing real-world data during training. Furthermore, our approach achieves zero-shot generalization in both action following and reconstruction tasks, thanks to its intrinsic geometric modeling. Notably, even without real-world data, its reconstruction performance is comparable with or even better than that of domain-specific models. Additionally, Aether employs camera trajectories as geometry-informed action spaces, enabling effective action-conditioned prediction and visual planning. We hope our work inspires the community to explore new frontiers in physically-reasonable world modeling and its applications. |
| title | Aether: Geometric-Aware Unified World Modeling |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Robotics |
| url | https://arxiv.org/abs/2503.18945 |