AirScape: An Aerial Generative World Model with Motion Controllability

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Baining, Tang, Rongze, Jia, Mingyuan, Wang, Ziyou, Man, Fanghang, Zhang, Xin, Shang, Yu, Zhang, Weichen, Wu, Wei, Gao, Chen, Chen, Xinlei, Li, Yong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915543064248320
author Zhao, Baining
Tang, Rongze
Jia, Mingyuan
Wang, Ziyou
Man, Fanghang
Zhang, Xin
Shang, Yu
Zhang, Weichen
Wu, Wei
Gao, Chen
Chen, Xinlei
Li, Yong
author_facet Zhao, Baining
Tang, Rongze
Jia, Mingyuan
Wang, Ziyou
Man, Fanghang
Zhang, Xin
Shang, Yu
Zhang, Weichen
Wu, Wei
Gao, Chen
Chen, Xinlei
Li, Yong
contents How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the first world model designed for six-degree-of-freedom aerial agents. AirScape predicts future observation sequences based on current visual inputs and motion intentions. Specifically, we construct a dataset for aerial world model training and testing, which consists of 11k video-intention pairs. This dataset includes first-person-view videos capturing diverse drone actions across a wide range of scenarios, with over 1,000 hours spent annotating the corresponding motion intentions. Then we develop a two-phase schedule to train a foundation model--initially devoid of embodied spatial knowledge--into a world model that is controllable by motion intentions and adheres to physical spatio-temporal constraints. Experimental results demonstrate that AirScape significantly outperforms existing foundation models in 3D spatial imagination capabilities, especially with over a 50% improvement in metrics reflecting motion alignment. The project is available at: https://embodiedcity.github.io/AirScape/.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08885
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AirScape: An Aerial Generative World Model with Motion Controllability
Zhao, Baining
Tang, Rongze
Jia, Mingyuan
Wang, Ziyou
Man, Fanghang
Zhang, Xin
Shang, Yu
Zhang, Weichen
Wu, Wei
Gao, Chen
Chen, Xinlei
Li, Yong
Robotics
Artificial Intelligence
How to enable agents to predict the outcomes of their own motion intentions in three-dimensional space has been a fundamental problem in embodied intelligence. To explore general spatial imagination capability, we present AirScape, the first world model designed for six-degree-of-freedom aerial agents. AirScape predicts future observation sequences based on current visual inputs and motion intentions. Specifically, we construct a dataset for aerial world model training and testing, which consists of 11k video-intention pairs. This dataset includes first-person-view videos capturing diverse drone actions across a wide range of scenarios, with over 1,000 hours spent annotating the corresponding motion intentions. Then we develop a two-phase schedule to train a foundation model--initially devoid of embodied spatial knowledge--into a world model that is controllable by motion intentions and adheres to physical spatio-temporal constraints. Experimental results demonstrate that AirScape significantly outperforms existing foundation models in 3D spatial imagination capabilities, especially with over a 50% improvement in metrics reflecting motion alignment. The project is available at: https://embodiedcity.github.io/AirScape/.
title AirScape: An Aerial Generative World Model with Motion Controllability
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2507.08885