PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Mao, Qing, Huang, Tianxin, Zhu, Yu, Sun, Jinqiu, Zhang, Yanning, Lee, Gim Hee
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914107961114624
author Mao, Qing
Huang, Tianxin
Zhu, Yu
Sun, Jinqiu
Zhang, Yanning
Lee, Gim Hee
author_facet Mao, Qing
Huang, Tianxin
Zhu, Yu
Sun, Jinqiu
Zhang, Yanning
Lee, Gim Hee
contents Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video interpolation and selecting key frames via a self-consistency score. However, the generated frames are often blurry due to small overlap inputs, and the selection strategies are slow and not explicitly aligned with pose estimation. To solve these cases, we propose Hybrid Video Generation (HVG) to synthesize clearer intermediate frames by coupling a video interpolation model with a pose-conditioned novel view synthesis model, where we also propose a Feature Matching Selector (FMS) based on feature correspondence to select intermediate frames appropriate for pose estimation from the synthesized results. Extensive experiments on Cambridge Landmarks, ScanNet, DL3DV-10K, and NAVI demonstrate that, compared to existing SOTA methods, PoseCrafter can obviously enhance the pose estimation performances, especially on examples with small or no overlap.
format Preprint
id arxiv_https___arxiv_org_abs_2510_19527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
Mao, Qing
Huang, Tianxin
Zhu, Yu
Sun, Jinqiu
Zhang, Yanning
Lee, Gim Hee
Computer Vision and Pattern Recognition
Pairwise camera pose estimation from sparsely overlapping image pairs remains a critical and unsolved challenge in 3D vision. Most existing methods struggle with image pairs that have small or no overlap. Recent approaches attempt to address this by synthesizing intermediate frames using video interpolation and selecting key frames via a self-consistency score. However, the generated frames are often blurry due to small overlap inputs, and the selection strategies are slow and not explicitly aligned with pose estimation. To solve these cases, we propose Hybrid Video Generation (HVG) to synthesize clearer intermediate frames by coupling a video interpolation model with a pose-conditioned novel view synthesis model, where we also propose a Feature Matching Selector (FMS) based on feature correspondence to select intermediate frames appropriate for pose estimation from the synthesized results. Extensive experiments on Cambridge Landmarks, ScanNet, DL3DV-10K, and NAVI demonstrate that, compared to existing SOTA methods, PoseCrafter can obviously enhance the pose estimation performances, especially on examples with small or no overlap.
title PoseCrafter: Extreme Pose Estimation with Hybrid Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.19527