Aether: Geometric-Aware Unified World Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aether Team, Zhu, Haoyi, Wang, Yifan, Zhou, Jianjun, Chang, Wenzheng, Zhou, Yang, Li, Zizun, Chen, Junyi, Shen, Chunhua, Pang, Jiangmiao, He, Tong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912503550705664
author Aether Team
Zhu, Haoyi
Wang, Yifan
Zhou, Jianjun
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Shen, Chunhua
Pang, Jiangmiao
He, Tong
author_facet Aether Team
Zhu, Haoyi
Wang, Yifan
Zhou, Jianjun
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Shen, Chunhua
Pang, Jiangmiao
He, Tong
contents The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware reasoning in world models by jointly optimizing three core capabilities: (1) 4D dynamic reconstruction, (2) action-conditioned video prediction, and (3) goal-conditioned visual planning. Through task-interleaved feature learning, Aether achieves synergistic knowledge sharing across reconstruction, prediction, and planning objectives. Building upon video generation models, our framework demonstrates zero-shot synthetic-to-real generalization despite never observing real-world data during training. Furthermore, our approach achieves zero-shot generalization in both action following and reconstruction tasks, thanks to its intrinsic geometric modeling. Notably, even without real-world data, its reconstruction performance is comparable with or even better than that of domain-specific models. Additionally, Aether employs camera trajectories as geometry-informed action spaces, enabling effective action-conditioned prediction and visual planning. We hope our work inspires the community to explore new frontiers in physically-reasonable world modeling and its applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18945
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aether: Geometric-Aware Unified World Modeling
Aether Team
Zhu, Haoyi
Wang, Yifan
Zhou, Jianjun
Chang, Wenzheng
Zhou, Yang
Li, Zizun
Chen, Junyi
Shen, Chunhua
Pang, Jiangmiao
He, Tong
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware reasoning in world models by jointly optimizing three core capabilities: (1) 4D dynamic reconstruction, (2) action-conditioned video prediction, and (3) goal-conditioned visual planning. Through task-interleaved feature learning, Aether achieves synergistic knowledge sharing across reconstruction, prediction, and planning objectives. Building upon video generation models, our framework demonstrates zero-shot synthetic-to-real generalization despite never observing real-world data during training. Furthermore, our approach achieves zero-shot generalization in both action following and reconstruction tasks, thanks to its intrinsic geometric modeling. Notably, even without real-world data, its reconstruction performance is comparable with or even better than that of domain-specific models. Additionally, Aether employs camera trajectories as geometry-informed action spaces, enabling effective action-conditioned prediction and visual planning. We hope our work inspires the community to explore new frontiers in physically-reasonable world modeling and its applications.
title Aether: Geometric-Aware Unified World Modeling
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2503.18945