AutoScape: Geometry-Consistent Long-Horizon Scene Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914109521395712 |
|---|---|
| author | Chen, Jiacheng Jiang, Ziyu Liang, Mingfu Zhuang, Bingbing Su, Jong-Chyi Garg, Sparsh Wu, Ying Chandraker, Manmohan |
| author_facet | Chen, Jiacheng Jiang, Ziyu Liang, Mingfu Zhuang, Bingbing Su, Jong-Chyi Garg, Sparsh Wu, Ying Chandraker, Manmohan |
| contents | This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_20726 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | AutoScape: Geometry-Consistent Long-Horizon Scene Generation Chen, Jiacheng Jiang, Ziyu Liang, Mingfu Zhuang, Bingbing Su, Jong-Chyi Garg, Sparsh Wu, Ying Chandraker, Manmohan Computer Vision and Pattern Recognition This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the scene's appearance and geometry. To maintain long-range geometric consistency, the model 1) jointly handles image and depth in a shared latent space, 2) explicitly conditions on the existing scene geometry (i.e., rendered point clouds) from previously generated keyframes, and 3) steers the sampling process with a warp-consistent guidance. Given high-quality RGB-D keyframes, a video diffusion model then interpolates between them to produce dense and coherent video frames. AutoScape generates realistic and geometrically consistent driving videos of over 20 seconds, improving the long-horizon FID and FVD scores over the prior state-of-the-art by 48.6\% and 43.0\%, respectively. |
| title | AutoScape: Geometry-Consistent Long-Horizon Scene Generation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2510.20726 |