Masked Depth Modeling for Spatial Perception

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Tan, Bin, Sun, Changjiang, Qin, Xiage, Adai, Hanat, Fu, Zelin, Zhou, Tianxiang, Zhang, Han, Xu, Yinghao, Zhu, Xing, Shen, Yujun, Xue, Nan
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918304776454144
author Tan, Bin
Sun, Changjiang
Qin, Xiage
Adai, Hanat
Fu, Zelin
Zhou, Tianxiang
Zhang, Han
Xu, Yinghao
Zhu, Xing
Shen, Yujun
Xue, Nan
author_facet Tan, Bin
Sun, Changjiang
Qin, Xiage
Adai, Hanat
Fu, Zelin
Zhou, Tianxiang
Zhang, Han
Xu, Yinghao
Zhu, Xing
Shen, Yujun
Xue, Nan
contents Spatial visual perception is a fundamental requirement in physical-world applications like autonomous driving and robotic manipulation, driven by the need to interact with 3D environments. Capturing pixel-aligned metric depth using RGB-D cameras would be the most viable way, yet it usually faces obstacles posed by hardware limitations and challenging imaging conditions, especially in the presence of specular or texture-less surfaces. In this work, we argue that the inaccuracies from depth sensors can be viewed as "masked" signals that inherently reflect underlying geometric ambiguities. Building on this motivation, we present LingBot-Depth, a depth completion model which leverages visual context to refine depth maps through masked depth modeling and incorporates an automated data curation pipeline for scalable training. It is encouraging to see that our model outperforms top-tier RGB-D cameras in terms of both depth precision and pixel coverage. Experimental results on a range of downstream tasks further suggest that LingBot-Depth offers an aligned latent representation across RGB and depth modalities. We release the code, checkpoint, and 3M RGB-depth pairs (including 2M real data and 1M simulated data) to the community of spatial perception.
format Preprint
id arxiv_https___arxiv_org_abs_2601_17895
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Masked Depth Modeling for Spatial Perception
Tan, Bin
Sun, Changjiang
Qin, Xiage
Adai, Hanat
Fu, Zelin
Zhou, Tianxiang
Zhang, Han
Xu, Yinghao
Zhu, Xing
Shen, Yujun
Xue, Nan
Computer Vision and Pattern Recognition
Robotics
Spatial visual perception is a fundamental requirement in physical-world applications like autonomous driving and robotic manipulation, driven by the need to interact with 3D environments. Capturing pixel-aligned metric depth using RGB-D cameras would be the most viable way, yet it usually faces obstacles posed by hardware limitations and challenging imaging conditions, especially in the presence of specular or texture-less surfaces. In this work, we argue that the inaccuracies from depth sensors can be viewed as "masked" signals that inherently reflect underlying geometric ambiguities. Building on this motivation, we present LingBot-Depth, a depth completion model which leverages visual context to refine depth maps through masked depth modeling and incorporates an automated data curation pipeline for scalable training. It is encouraging to see that our model outperforms top-tier RGB-D cameras in terms of both depth precision and pixel coverage. Experimental results on a range of downstream tasks further suggest that LingBot-Depth offers an aligned latent representation across RGB and depth modalities. We release the code, checkpoint, and 3M RGB-depth pairs (including 2M real data and 1M simulated data) to the community of spatial perception.
title Masked Depth Modeling for Spatial Perception
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2601.17895