A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dong, Yitong, Li, Yijin, Huang, Zhaoyang, Bian, Weikang, Liu, Jingbo, Bao, Hujun, Cui, Zhaopeng, Li, Hongsheng, Zhang, Guofeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916513939718144
author Dong, Yitong
Li, Yijin
Huang, Zhaoyang
Bian, Weikang
Liu, Jingbo
Bao, Hujun
Cui, Zhaopeng
Li, Hongsheng
Zhang, Guofeng
author_facet Dong, Yitong
Li, Yijin
Huang, Zhaoyang
Bian, Weikang
Liu, Jingbo
Bao, Hujun
Cui, Zhaopeng
Li, Hongsheng
Zhang, Guofeng
contents In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks&Temple benchmark demonstrate the effectiveness of our method. The code is available at our project page: https://zju3dv.github.io/GD-PoseMVS/.
format Preprint
id arxiv_https___arxiv_org_abs_2411_01893
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding
Dong, Yitong
Li, Yijin
Huang, Zhaoyang
Bian, Weikang
Liu, Jingbo
Bao, Hujun
Cui, Zhaopeng
Li, Hongsheng
Zhang, Guofeng
Computer Vision and Pattern Recognition
In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images. Specifically, we introduce a Multi-view Disparity Attention (MDA) module to aggregate long-range context information within and across multi-view images. Considering the asymmetry of the epipolar disparity flow, the key to our method lies in accurately modeling multi-view geometric constraints. We integrate pose embedding to encapsulate information such as multi-view camera poses, providing implicit geometric constraints for multi-view disparity feature fusion dominated by attention. Additionally, we construct corresponding hidden states for each source image due to significant differences in the observation quality of the same pixel in the reference frame across multiple source frames. We explicitly estimate the quality of the current pixel corresponding to sampled points on the epipolar line of the source image and dynamically update hidden states through the uncertainty estimation module. Extensive results on the DTU dataset and Tanks&Temple benchmark demonstrate the effectiveness of our method. The code is available at our project page: https://zju3dv.github.io/GD-PoseMVS/.
title A Global Depth-Range-Free Multi-View Stereo Transformer Network with Pose Embedding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.01893