Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Gangwei, Lin, Haotong, Luo, Hongcheng, Wang, Xianqi, Yao, Jingfeng, Zhu, Lianghui, Pu, Yuechuan, Chi, Cheng, Sun, Haiyang, Wang, Bing, Chen, Guang, Ye, Hangjun, Peng, Sida, Yang, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026)
by: Xu, Gangwei, et al.
Published: (2026)
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026)
by: Chi, Cheng, et al.
Published: (2026)
BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
Uni-Gaussians: Unifying Camera and Lidar Simulation with Gaussians for Dynamic Driving Scenarios
by: Yuan, Zikang, et al.
Published: (2025)
by: Yuan, Zikang, et al.
Published: (2025)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)
by: Wang, Shuyun, et al.
Published: (2025)
PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
by: Wang, Xianqi, et al.
Published: (2026)
by: Wang, Xianqi, et al.
Published: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
Selective-Stereo: Adaptive Frequency Information Selection for Stereo Matching
by: Wang, Xianqi, et al.
Published: (2024)
by: Wang, Xianqi, et al.
Published: (2024)
ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025)
by: Xia, Tianze, et al.
Published: (2025)
FlowMamba: Learning Point Cloud Scene Flow with Global Motion Propagation
by: Lin, Min, et al.
Published: (2024)
by: Lin, Min, et al.
Published: (2024)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
by: Lu, Wenquan, et al.
Published: (2025)
by: Lu, Wenquan, et al.
Published: (2025)
Generalized Geometry Encoding Volume for Real-time Stereo Matching
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
PCSTracker: Long-Term Scene Flow Estimation for Point Cloud Sequences
by: Lin, Min, et al.
Published: (2026)
by: Lin, Min, et al.
Published: (2026)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026)
by: Tan, Kaiyuan, et al.
Published: (2026)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
Toward Physically Consistent Driving Video World Models under Challenging Trajectories
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
FasterDiT: Towards Faster Diffusion Transformers Training without Architecture Modification
by: Yao, Jingfeng, et al.
Published: (2024)
by: Yao, Jingfeng, et al.
Published: (2024)
InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
by: Chen, Xiaoxue, et al.
Published: (2025)
by: Chen, Xiaoxue, et al.
Published: (2025)
IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching
by: Xu, Gangwei, et al.
Published: (2024)
by: Xu, Gangwei, et al.
Published: (2024)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Precise upper deviation estimates for the maximum of a branching random walk
by: Luo, Lianghui
Published: (2024)
by: Luo, Lianghui
Published: (2024)
The extremal process of two-speed branching random walk
by: Luo, Lianghui
Published: (2025)
by: Luo, Lianghui
Published: (2025)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Leveraging Stable Diffusion for Monocular Depth Estimation via Image Semantic Encoding
by: Xia, Jingming, et al.
Published: (2025)
by: Xia, Jingming, et al.
Published: (2025)
Laytrol: Preserving Pretrained Knowledge in Layout Control for Multimodal Diffusion Transformers
by: Huang, Sida, et al.
Published: (2025)
by: Huang, Sida, et al.
Published: (2025)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
by: Yao, Jingfeng, et al.
Published: (2025)
by: Yao, Jingfeng, et al.
Published: (2025)
MLE-based Device Activity Detection under Rician Fading for Massive Grant-free Access with Perfect and Imperfect Synchronization
by: Liu, Wang, et al.
Published: (2023)
by: Liu, Wang, et al.
Published: (2023)
Mitigate Negative Transfer with Similarity Heuristic Lifelong Prompt Tuning
by: Wu, Chenyuan, et al.
Published: (2024)
by: Wu, Chenyuan, et al.
Published: (2024)
Xiaomi Auto World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving
by: Zhou, Lijun, et al.
Published: (2026)
by: Zhou, Lijun, et al.
Published: (2026)
Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
Upper moderate deviation probabilities for the maximum of a branching random walk
by: Chataignier, Louis, et al.
Published: (2026)
by: Chataignier, Louis, et al.
Published: (2026)
Upper deviation probabilities for level sets of a supercritical branching random walk
by: Zhang, Shuxiong, et al.
Published: (2024)
by: Zhang, Shuxiong, et al.
Published: (2024)
WeakTr: Exploring Plain Vision Transformer for Weakly-supervised Semantic Segmentation
by: Zhu, Lianghui, et al.
Published: (2023)
by: Zhu, Lianghui, et al.
Published: (2023)
BANet: Bilateral Aggregation Network for Mobile Stereo Matching
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
by: Yao, Yuxuan, et al.
Published: (2026)
by: Yao, Yuxuan, et al.
Published: (2026)
Similar Items
-
Pixel-Perfect Visual Geometry Estimation
by: Xu, Gangwei, et al.
Published: (2026) -
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026) -
BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation
by: Xu, Gangwei, et al.
Published: (2025) -
Uni-Gaussians: Unifying Camera and Lidar Simulation with Gaussians for Dynamic Driving Scenarios
by: Yuan, Zikang, et al.
Published: (2025) -
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)