PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xia, Yifei, Weng, Shuchen, Yang, Siqi, Liu, Jingqi, Zhu, Chengxuan, Teng, Minggui, Jia, Zijian, Jiang, Han, Shi, Boxin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909659618607104
author Xia, Yifei
Weng, Shuchen
Yang, Siqi
Liu, Jingqi
Zhu, Chengxuan
Teng, Minggui
Jia, Zijian
Jiang, Han
Shi, Boxin
author_facet Xia, Yifei
Weng, Shuchen
Yang, Siqi
Liu, Jingqi
Zhu, Chengxuan
Teng, Minggui
Jia, Zijian
Jiang, Han
Shi, Boxin
contents Panoramic video generation enables immersive 360° content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks. Our project page is available at https://panowan.variantconst.com.
format Preprint
id arxiv_https___arxiv_org_abs_2505_22016
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
Xia, Yifei
Weng, Shuchen
Yang, Siqi
Liu, Jingqi
Zhu, Chengxuan
Teng, Minggui
Jia, Zijian
Jiang, Han
Shi, Boxin
Computer Vision and Pattern Recognition
Panoramic video generation enables immersive 360° content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks. Our project page is available at https://panowan.variantconst.com.
title PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.22016