Scalable Autoregressive Monocular Depth Estimation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Jinhong, Liu, Jian, Tang, Dongqi, Wang, Weiqiang, Li, Wentong, Chen, Danny, Chen, Jintai, Wu, Jian
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917960583479296
author Wang, Jinhong
Liu, Jian
Tang, Dongqi
Wang, Weiqiang
Li, Wentong
Chen, Danny
Chen, Jintai
Wu, Jian
author_facet Wang, Jinhong
Liu, Jian
Tang, Dongqi
Wang, Weiqiang
Li, Wentong
Chen, Danny
Chen, Jintai
Wu, Jian
contents This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise casual mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-of-the-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11361
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Scalable Autoregressive Monocular Depth Estimation
Wang, Jinhong
Liu, Jian
Tang, Dongqi
Wang, Weiqiang
Li, Wentong
Chen, Danny
Chen, Jintai
Wu, Jian
Computer Vision and Pattern Recognition
This paper shows that the autoregressive model is an effective and scalable monocular depth estimator. Our idea is simple: We tackle the monocular depth estimation (MDE) task with an autoregressive prediction paradigm, based on two core designs. First, our depth autoregressive model (DAR) treats the depth map of different resolutions as a set of tokens, and conducts the low-to-high resolution autoregressive objective with a patch-wise casual mask. Second, our DAR recursively discretizes the entire depth range into more compact intervals, and attains the coarse-to-fine granularity autoregressive objective in an ordinal-regression manner. By coupling these two autoregressive objectives, our DAR establishes new state-of-the-art (SOTA) on KITTI and NYU Depth v2 by clear margins. Further, our scalable approach allows us to scale the model up to 2.0B and achieve the best RMSE of 1.799 on the KITTI dataset (5% improvement) compared to 1.896 by the current SOTA (Depth Anything). DAR further showcases zero-shot generalization ability on unseen datasets. These results suggest that DAR yields superior performance with an autoregressive prediction paradigm, providing a promising approach to equip modern autoregressive large models (e.g., GPT-4o) with depth estimation capabilities.
title Scalable Autoregressive Monocular Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.11361