SE360: Semantic Edit in 360$^\circ$ Panoramas via Hierarchical Data Construction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhong, Haoyi, Zhang, Fang-Lue, Chalmers, Andrew, Rhee, Taehyun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912785086021632
author Zhong, Haoyi
Zhang, Fang-Lue
Chalmers, Andrew
Rhee, Taehyun
author_facet Zhong, Haoyi
Zhang, Fang-Lue
Chalmers, Andrew
Rhee, Taehyun
contents While instruction-based image editing is emerging, extending it to 360$^\circ$ panoramas introduces additional challenges. Existing methods often produce implausible results in both equirectangular projections (ERP) and perspective views. To address these limitations, we propose SE360, a novel framework for multi-condition guided object editing in 360$^\circ$ panoramas. At its core is a novel coarse-to-fine autonomous data generation pipeline without manual intervention. This pipeline leverages a Vision-Language Model (VLM) and adaptive projection adjustment for hierarchical analysis, ensuring the holistic segmentation of objects and their physical context. The resulting data pairs are both semantically meaningful and geometrically consistent, even when sourced from unlabeled panoramas. Furthermore, we introduce a cost-effective, two-stage data refinement strategy to improve data realism and mitigate model overfitting to erase artifacts. Based on the constructed dataset, we train a Transformer-based diffusion model to allow flexible object editing guided by text, mask, or reference image in 360$^\circ$ panoramas. Our experiments demonstrate that our method outperforms existing methods in both visual quality and semantic accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2512_19943
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SE360: Semantic Edit in 360$^\circ$ Panoramas via Hierarchical Data Construction
Zhong, Haoyi
Zhang, Fang-Lue
Chalmers, Andrew
Rhee, Taehyun
Computer Vision and Pattern Recognition
While instruction-based image editing is emerging, extending it to 360$^\circ$ panoramas introduces additional challenges. Existing methods often produce implausible results in both equirectangular projections (ERP) and perspective views. To address these limitations, we propose SE360, a novel framework for multi-condition guided object editing in 360$^\circ$ panoramas. At its core is a novel coarse-to-fine autonomous data generation pipeline without manual intervention. This pipeline leverages a Vision-Language Model (VLM) and adaptive projection adjustment for hierarchical analysis, ensuring the holistic segmentation of objects and their physical context. The resulting data pairs are both semantically meaningful and geometrically consistent, even when sourced from unlabeled panoramas. Furthermore, we introduce a cost-effective, two-stage data refinement strategy to improve data realism and mitigate model overfitting to erase artifacts. Based on the constructed dataset, we train a Transformer-based diffusion model to allow flexible object editing guided by text, mask, or reference image in 360$^\circ$ panoramas. Our experiments demonstrate that our method outperforms existing methods in both visual quality and semantic accuracy.
title SE360: Semantic Edit in 360$^\circ$ Panoramas via Hierarchical Data Construction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.19943