Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lin, Haotong, Peng, Sida, Chen, Jingxiao, Peng, Songyou, Sun, Jiaming, Liu, Minghuan, Bao, Hujun, Feng, Jiashi, Zhou, Xiaowei, Kang, Bingyi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917365010137088
author Lin, Haotong
Peng, Sida
Chen, Jingxiao
Peng, Songyou
Sun, Jiaming
Liu, Minghuan
Bao, Hujun
Feng, Jiashi
Zhou, Xiaowei
Kang, Bingyi
author_facet Lin, Haotong
Peng, Sida
Chen, Jingxiao
Peng, Songyou
Sun, Jiaming
Liu, Minghuan
Bao, Hujun
Feng, Jiashi
Zhou, Xiaowei
Kang, Bingyi
contents Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost LiDAR as the prompt to guide the Depth Anything model for accurate metric depth output, achieving up to 4K resolution. Our approach centers on a concise prompt fusion design that integrates the LiDAR at multiple scales within the depth decoder. To address training challenges posed by limited datasets containin both LiDAR depth and precise GT depth, we propose a scalable data pipeline that includes synthetic data LiDAR simulation and real data pseudo GT depth generation. Our approach sets new state-of-the-arts on the ARKitScenes and ScanNet++ datasets and benefits downstream applications, including 3D reconstruction and generalized robotic grasping.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14015
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
Lin, Haotong
Peng, Sida
Chen, Jingxiao
Peng, Songyou
Sun, Jiaming
Liu, Minghuan
Bao, Hujun
Feng, Jiashi
Zhou, Xiaowei
Kang, Bingyi
Computer Vision and Pattern Recognition
Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost LiDAR as the prompt to guide the Depth Anything model for accurate metric depth output, achieving up to 4K resolution. Our approach centers on a concise prompt fusion design that integrates the LiDAR at multiple scales within the depth decoder. To address training challenges posed by limited datasets containin both LiDAR depth and precise GT depth, we propose a scalable data pipeline that includes synthetic data LiDAR simulation and real data pseudo GT depth generation. Our approach sets new state-of-the-arts on the ARKitScenes and ScanNet++ datasets and benefits downstream applications, including 3D reconstruction and generalized robotic grasping.
title Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.14015