Pixel-Perfect Visual Geometry Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Gangwei, Lin, Haotong, Luo, Hongcheng, Sun, Haiyang, Wang, Bing, Chen, Guang, Peng, Sida, Ye, Hangjun, Yang, Xin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation
by: Xu, Gangwei, et al.
Published: (2025)
by: Xu, Gangwei, et al.
Published: (2025)
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026)
by: Chi, Cheng, et al.
Published: (2026)
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)
by: Wang, Shuyun, et al.
Published: (2025)
ExtraGS: Geometric-Aware Trajectory Extrapolation with Uncertainty-Guided Generative Priors
by: Tan, Kaiyuan, et al.
Published: (2025)
by: Tan, Kaiyuan, et al.
Published: (2025)
Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
by: Guo, Xiangyu, et al.
Published: (2025)
by: Guo, Xiangyu, et al.
Published: (2025)
PCSTracker: Long-Term Scene Flow Estimation for Point Cloud Sequences
by: Lin, Min, et al.
Published: (2026)
by: Lin, Min, et al.
Published: (2026)
UFO: Unifying Feed-Forward and Optimization-based Methods for Large Driving Scene Modeling
by: Tan, Kaiyuan, et al.
Published: (2026)
by: Tan, Kaiyuan, et al.
Published: (2026)
Toward Physically Consistent Driving Video World Models under Challenging Trajectories
by: Zhou, Jiawei, et al.
Published: (2026)
by: Zhou, Jiawei, et al.
Published: (2026)
DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
by: Chen, Xiaoxue, et al.
Published: (2025)
by: Chen, Xiaoxue, et al.
Published: (2025)
Generalized Geometry Encoding Volume for Real-time Stereo Matching
by: Liu, Jiaxin, et al.
Published: (2025)
by: Liu, Jiaxin, et al.
Published: (2025)
ADGaussian: Generalizable Gaussian Splatting for Autonomous Driving via Multi-modal Joint Learning
by: Song, Qi, et al.
Published: (2025)
by: Song, Qi, et al.
Published: (2025)
ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
by: Li, Yongkang, et al.
Published: (2025)
by: Li, Yongkang, et al.
Published: (2025)
Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
by: Lin, Haotong, et al.
Published: (2024)
by: Lin, Haotong, et al.
Published: (2024)
IGEV++: Iterative Multi-range Geometry Encoding Volumes for Stereo Matching
by: Xu, Gangwei, et al.
Published: (2024)
by: Xu, Gangwei, et al.
Published: (2024)
Street Gaussians: Modeling Dynamic Urban Scenes with Gaussian Splatting
by: Yan, Yunzhi, et al.
Published: (2024)
by: Yan, Yunzhi, et al.
Published: (2024)
WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
by: Zhu, Ziyue, et al.
Published: (2025)
by: Zhu, Ziyue, et al.
Published: (2025)
LVDrive: Latent Visual Representation Enhanced Vision-Language-Action Autonomous Driving Model
by: Mei, Xiaodong, et al.
Published: (2026)
by: Mei, Xiaodong, et al.
Published: (2026)
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
by: Zhang, Zhaoxing, et al.
Published: (2024)
by: Zhang, Zhaoxing, et al.
Published: (2024)
Multi-view Reconstruction via SfM-guided Monocular Depth Estimation
by: Guo, Haoyu, et al.
Published: (2025)
by: Guo, Haoyu, et al.
Published: (2025)
FlowMamba: Learning Point Cloud Scene Flow with Global Motion Propagation
by: Lin, Min, et al.
Published: (2024)
by: Lin, Min, et al.
Published: (2024)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
by: Zheng, Haitian, et al.
Published: (2025)
by: Zheng, Haitian, et al.
Published: (2025)
Selective-Stereo: Adaptive Frequency Information Selection for Stereo Matching
by: Wang, Xianqi, et al.
Published: (2024)
by: Wang, Xianqi, et al.
Published: (2024)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
DriveLaW:Unifying Planning and Video Generation in a Latent Driving World
by: Xia, Tianze, et al.
Published: (2025)
by: Xia, Tianze, et al.
Published: (2025)
Memory-Efficient Optical Flow via Radius-Distribution Orthogonal Cost Volume
by: Xu, Gangwei, et al.
Published: (2023)
by: Xu, Gangwei, et al.
Published: (2023)
InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Painting 3D Nature in 2D: View Synthesis of Natural Scenes from a Single Semantic Mask
by: Zhang, Shangzan, et al.
Published: (2023)
by: Zhang, Shangzan, et al.
Published: (2023)
LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving
by: Luo, Yuechen, et al.
Published: (2026)
by: Luo, Yuechen, et al.
Published: (2026)
Hybrid Cost Volume for Memory-Efficient Optical Flow
by: Zhao, Yang, et al.
Published: (2024)
by: Zhao, Yang, et al.
Published: (2024)
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
by: Xu, Guangkai, et al.
Published: (2026)
by: Xu, Guangkai, et al.
Published: (2026)
HDRFlow: Real-Time HDR Video Reconstruction with Large Motions
by: Xu, Gangwei, et al.
Published: (2024)
by: Xu, Gangwei, et al.
Published: (2024)
Depth Anything 3: Recovering the Visual Space from Any Views
by: Lin, Haotong, et al.
Published: (2025)
by: Lin, Haotong, et al.
Published: (2025)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
by: Wang, Xianqi, et al.
Published: (2026)
by: Wang, Xianqi, et al.
Published: (2026)
SpatialTracker: Tracking Any 2D Pixels in 3D Space
by: Xiao, Yuxi, et al.
Published: (2024)
by: Xiao, Yuxi, et al.
Published: (2024)
ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous Parking
by: Wei, Xiaobao, et al.
Published: (2026)
by: Wei, Xiaobao, et al.
Published: (2026)
From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
by: Liu, Yepeng, et al.
Published: (2026)
by: Liu, Yepeng, et al.
Published: (2026)
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
by: Xu, Zhen, et al.
Published: (2025)
by: Xu, Zhen, et al.
Published: (2025)
Similar Items
-
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
by: Xu, Gangwei, et al.
Published: (2025) -
BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal Correlation
by: Xu, Gangwei, et al.
Published: (2025) -
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
by: Chi, Cheng, et al.
Published: (2026) -
ViSE: A Systematic Approach to Vision-Only Street-View Extrapolation
by: Tan, Kaiyuan, et al.
Published: (2025) -
Mirage: One-Step Video Diffusion for Photorealistic and Coherent Asset Editing in Driving Scenes
by: Wang, Shuyun, et al.
Published: (2025)