SGFormer: Spherical Geometry Transformer for 360 Depth Estimation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Junsong, Chen, Zisong, Lin, Chunyu, Nie, Lang, Shen, Zhijie, Liao, Kang, Huang, Junda, Zhao, Yao
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929730531360768
author Zhang, Junsong
Chen, Zisong
Lin, Chunyu
Nie, Lang
Shen, Zhijie
Liao, Kang
Huang, Junda
Zhao, Yao
author_facet Zhang, Junsong
Chen, Zisong
Lin, Chunyu
Nie, Lang
Shen, Zhijie
Liao, Kang
Huang, Junda
Zhao, Yao
contents Panoramic distortion poses a significant challenge in 360 depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, which can result in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar re-projection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions.
format Preprint
id arxiv_https___arxiv_org_abs_2404_14979
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SGFormer: Spherical Geometry Transformer for 360 Depth Estimation
Zhang, Junsong
Chen, Zisong
Lin, Chunyu
Nie, Lang
Shen, Zhijie
Liao, Kang
Huang, Junda
Zhao, Yao
Computer Vision and Pattern Recognition
Artificial Intelligence
Panoramic distortion poses a significant challenge in 360 depth estimation, particularly pronounced at the north and south poles. Existing methods either adopt a bi-projection fusion strategy to remove distortions or model long-range dependencies to capture global structures, which can result in either unclear structure or insufficient local perception. In this paper, we propose a spherical geometry transformer, named SGFormer, to address the above issues, with an innovative step to integrate spherical geometric priors into vision transformers. To this end, we retarget the transformer decoder to a spherical prior decoder (termed SPDecoder), which endeavors to uphold the integrity of spherical structures during decoding. Concretely, we leverage bipolar re-projection, circular rotation, and curve local embedding to preserve the spherical characteristics of equidistortion, continuity, and surface distance, respectively. Furthermore, we present a query-based global conditional position embedding to compensate for spatial structure at varying resolutions. It not only boosts the global perception of spatial position but also sharpens the depth structure across different patches. Finally, we conduct extensive experiments on popular benchmarks, demonstrating our superiority over state-of-the-art solutions.
title SGFormer: Spherical Geometry Transformer for 360 Depth Estimation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2404.14979