Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Gangwei, Lin, Haotong, Luo, Hongcheng, Wang, Xianqi, Yao, Jingfeng, Zhu, Lianghui, Pu, Yuechuan, Chi, Cheng, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Peng, Sida, Yang, Xin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911238818103296
author Xu, Gangwei
Lin, Haotong
Luo, Hongcheng
Wang, Xianqi
Yao, Jingfeng
Zhu, Lianghui
Pu, Yuechuan
Chi, Cheng
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Peng, Sida
Yang, Xin
author_facet Xu, Gangwei
Lin, Haotong
Luo, Hongcheng
Wang, Xianqi
Yao, Jingfeng
Zhu, Lianghui
Pu, Yuechuan
Chi, Cheng
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Peng, Sida
Yang, Xin
contents This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation models fine-tune Stable Diffusion and achieve impressive performance. However, they require a VAE to compress depth maps into latent space, which inevitably introduces \textit{flying pixels} at edges and details. Our model addresses this challenge by directly performing diffusion generation in the pixel space, avoiding VAE-induced artifacts. To overcome the high complexity associated with pixel-space generation, we introduce two novel designs: 1) Semantics-Prompted Diffusion Transformers (SP-DiT), which incorporate semantic representations from vision foundation models into DiT to prompt the diffusion process, thereby preserving global semantic consistency while enhancing fine-grained visual details; and 2) Cascade DiT Design that progressively increases the number of tokens to further enhance efficiency and accuracy. Our model achieves the best performance among all published generative models across five benchmarks, and significantly outperforms all other models in edge-aware point cloud evaluation.
format Preprint
id arxiv_https___arxiv_org_abs_2510_07316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Xu, Gangwei
Lin, Haotong
Luo, Hongcheng
Wang, Xianqi
Yao, Jingfeng
Zhu, Lianghui
Pu, Yuechuan
Chi, Cheng
Sun, Haiyang
Wang, Bing
Chen, Guang
Ye, Hangjun
Peng, Sida
Yang, Xin
Computer Vision and Pattern Recognition
This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation models fine-tune Stable Diffusion and achieve impressive performance. However, they require a VAE to compress depth maps into latent space, which inevitably introduces \textit{flying pixels} at edges and details. Our model addresses this challenge by directly performing diffusion generation in the pixel space, avoiding VAE-induced artifacts. To overcome the high complexity associated with pixel-space generation, we introduce two novel designs: 1) Semantics-Prompted Diffusion Transformers (SP-DiT), which incorporate semantic representations from vision foundation models into DiT to prompt the diffusion process, thereby preserving global semantic consistency while enhancing fine-grained visual details; and 2) Cascade DiT Design that progressively increases the number of tokens to further enhance efficiency and accuracy. Our model achieves the best performance among all published generative models across five benchmarks, and significantly outperforms all other models in edge-aware point cloud evaluation.
title Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.07316