Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Wenyu, Liu, Sidun, Qiao, Peng, Dou, Yong, Hu, Tongrui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917098754670592
author Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
Hu, Tongrui
author_facet Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
Hu, Tongrui
contents We present Muskie, a native multi-view vision backbone designed for 3D vision tasks. Unlike existing models, which are frame-wise and exhibit limited multi-view consistency, Muskie is designed to process multiple views simultaneously and introduce multi-view consistency in pre-training stage. Muskie is trained to reconstruct heavily masked content in one view by finding and utilizing geometric correspondences from other views. Through this pretext task and our proposed aggressive masking strategy, the model implicitly to learn view-invariant features and develop strong geometric understanding without any 3D supervision. Compared with state-of-the-art frame-wise backbones such as DINO, Muskie achieves higher multi-view correspondence accuracy. Furthermore, we demonstrate that using Muskie as a backbone consistently enhances performance on downstream 3D tasks, including camera pose estimation and pointmap reconstruction. Codes are publicly available at https://leo-frank.github.io/Muskie/
format Preprint
id arxiv_https___arxiv_org_abs_2511_18115
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
Li, Wenyu
Liu, Sidun
Qiao, Peng
Dou, Yong
Hu, Tongrui
Computer Vision and Pattern Recognition
We present Muskie, a native multi-view vision backbone designed for 3D vision tasks. Unlike existing models, which are frame-wise and exhibit limited multi-view consistency, Muskie is designed to process multiple views simultaneously and introduce multi-view consistency in pre-training stage. Muskie is trained to reconstruct heavily masked content in one view by finding and utilizing geometric correspondences from other views. Through this pretext task and our proposed aggressive masking strategy, the model implicitly to learn view-invariant features and develop strong geometric understanding without any 3D supervision. Compared with state-of-the-art frame-wise backbones such as DINO, Muskie achieves higher multi-view correspondence accuracy. Furthermore, we demonstrate that using Muskie as a backbone consistently enhances performance on downstream 3D tasks, including camera pose estimation and pointmap reconstruction. Codes are publicly available at https://leo-frank.github.io/Muskie/
title Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.18115