DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Haoran, Zhang, Dizhe, Li, Xiangtai, Du, Bo, Qi, Lu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914090733010944
author Feng, Haoran
Zhang, Dizhe
Li, Xiangtai
Du, Bo
Qi, Lu
author_facet Feng, Haoran
Zhang, Dizhe
Li, Xiangtai
Du, Bo
Qi, Lu
contents In this work, we propose DiT360, a DiT-based framework that performs hybrid training on perspective and panoramic data for panoramic image generation. For the issues of maintaining geometric fidelity and photorealism in generation quality, we attribute the main reason to the lack of large-scale, high-quality, real-world panoramic data, where such a data-centric view differs from prior methods that focus on model design. Basically, DiT360 has several key modules for inter-domain transformation and intra-domain augmentation, applied at both the pre-VAE image level and the post-VAE token level. At the image level, we incorporate cross-domain knowledge through perspective image guidance and panoramic refinement, which enhance perceptual quality while regularizing diversity and photorealism. At the token level, hybrid supervision is applied across multiple modules, which include circular padding for boundary continuity, yaw loss for rotational robustness, and cube loss for distortion awareness. Extensive experiments on text-to-panorama, inpainting, and outpainting tasks demonstrate that our method achieves better boundary consistency and image fidelity across eleven quantitative metrics. Our code is available at https://github.com/Insta360-Research-Team/DiT360.
format Preprint
id arxiv_https___arxiv_org_abs_2510_11712
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
Feng, Haoran
Zhang, Dizhe
Li, Xiangtai
Du, Bo
Qi, Lu
Computer Vision and Pattern Recognition
In this work, we propose DiT360, a DiT-based framework that performs hybrid training on perspective and panoramic data for panoramic image generation. For the issues of maintaining geometric fidelity and photorealism in generation quality, we attribute the main reason to the lack of large-scale, high-quality, real-world panoramic data, where such a data-centric view differs from prior methods that focus on model design. Basically, DiT360 has several key modules for inter-domain transformation and intra-domain augmentation, applied at both the pre-VAE image level and the post-VAE token level. At the image level, we incorporate cross-domain knowledge through perspective image guidance and panoramic refinement, which enhance perceptual quality while regularizing diversity and photorealism. At the token level, hybrid supervision is applied across multiple modules, which include circular padding for boundary continuity, yaw loss for rotational robustness, and cube loss for distortion awareness. Extensive experiments on text-to-panorama, inpainting, and outpainting tasks demonstrate that our method achieves better boundary consistency and image fidelity across eleven quantitative metrics. Our code is available at https://github.com/Insta360-Research-Team/DiT360.
title DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.11712