CodingHomo: Bootstrapping Deep Homography With Video Coding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yike, Li, Haipeng, Liu, Shuaicheng, Zeng, Bing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910912766541824
author Liu, Yike
Li, Haipeng
Liu, Shuaicheng
Zeng, Bing
author_facet Liu, Yike
Li, Haipeng
Liu, Shuaicheng
Zeng, Bing
contents Homography estimation is a fundamental task in computer vision with applications in diverse fields. Recent advances in deep learning have improved homography estimation, particularly with unsupervised learning approaches, offering increased robustness and generalizability. However, accurately predicting homography, especially in complex motions, remains a challenge. In response, this work introduces a novel method leveraging video coding, particularly by harnessing inherent motion vectors (MVs) present in videos. We present CodingHomo, an unsupervised framework for homography estimation. Our framework features a Mask-Guided Fusion (MGF) module that identifies and utilizes beneficial features among the MVs, thereby enhancing the accuracy of homography prediction. Additionally, the Mask-Guided Homography Estimation (MGHE) module is presented for eliminating undesired features in the coarse-to-fine homography refinement process. CodingHomo outperforms existing state-of-the-art unsupervised methods, delivering good robustness and generalizability. The code and dataset are available at: \href{github}{https://github.com/liuyike422/CodingHomo
format Preprint
id arxiv_https___arxiv_org_abs_2504_12165
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CodingHomo: Bootstrapping Deep Homography With Video Coding
Liu, Yike
Li, Haipeng
Liu, Shuaicheng
Zeng, Bing
Computer Vision and Pattern Recognition
Homography estimation is a fundamental task in computer vision with applications in diverse fields. Recent advances in deep learning have improved homography estimation, particularly with unsupervised learning approaches, offering increased robustness and generalizability. However, accurately predicting homography, especially in complex motions, remains a challenge. In response, this work introduces a novel method leveraging video coding, particularly by harnessing inherent motion vectors (MVs) present in videos. We present CodingHomo, an unsupervised framework for homography estimation. Our framework features a Mask-Guided Fusion (MGF) module that identifies and utilizes beneficial features among the MVs, thereby enhancing the accuracy of homography prediction. Additionally, the Mask-Guided Homography Estimation (MGHE) module is presented for eliminating undesired features in the coarse-to-fine homography refinement process. CodingHomo outperforms existing state-of-the-art unsupervised methods, delivering good robustness and generalizability. The code and dataset are available at: \href{github}{https://github.com/liuyike422/CodingHomo
title CodingHomo: Bootstrapping Deep Homography With Video Coding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.12165