Sequence Matters: Harnessing Video Models in 3D Super-Resolution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ko, Hyun-kyu, Park, Dongheok, Park, Youngin, Lee, Byeonghyeon, Han, Juhee, Park, Eunbyung
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910758345900032
author Ko, Hyun-kyu
Park, Dongheok
Park, Youngin
Lee, Byeonghyeon
Han, Juhee
Park, Eunbyung
author_facet Ko, Hyun-kyu
Park, Dongheok
Park, Youngin
Lee, Byeonghyeon
Han, Juhee
Park, Eunbyung
contents 3D super-resolution aims to reconstruct high-fidelity 3D models from low-resolution (LR) multi-view images. Early studies primarily focused on single-image super-resolution (SISR) models to upsample LR images into high-resolution images. However, these methods often lack view consistency because they operate independently on each image. Although various post-processing techniques have been extensively explored to mitigate these inconsistencies, they have yet to fully resolve the issues. In this paper, we perform a comprehensive study of 3D super-resolution by leveraging video super-resolution (VSR) models. By utilizing VSR models, we ensure a higher degree of spatial consistency and can reference surrounding spatial information, leading to more accurate and detailed reconstructions. Our findings reveal that VSR models can perform remarkably well even on sequences that lack precise spatial alignment. Given this observation, we propose a simple yet practical approach to align LR images without involving fine-tuning or generating 'smooth' trajectory from the trained 3D models over LR images. The experimental results show that the surprisingly simple algorithms can achieve the state-of-the-art results of 3D super-resolution tasks on standard benchmark datasets, such as the NeRF-synthetic and MipNeRF-360 datasets. Project page: https://ko-lani.github.io/Sequence-Matters
format Preprint
id arxiv_https___arxiv_org_abs_2412_11525
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sequence Matters: Harnessing Video Models in 3D Super-Resolution
Ko, Hyun-kyu
Park, Dongheok
Park, Youngin
Lee, Byeonghyeon
Han, Juhee
Park, Eunbyung
Computer Vision and Pattern Recognition
68U10, 68T10
I.4.5; I.2.10
3D super-resolution aims to reconstruct high-fidelity 3D models from low-resolution (LR) multi-view images. Early studies primarily focused on single-image super-resolution (SISR) models to upsample LR images into high-resolution images. However, these methods often lack view consistency because they operate independently on each image. Although various post-processing techniques have been extensively explored to mitigate these inconsistencies, they have yet to fully resolve the issues. In this paper, we perform a comprehensive study of 3D super-resolution by leveraging video super-resolution (VSR) models. By utilizing VSR models, we ensure a higher degree of spatial consistency and can reference surrounding spatial information, leading to more accurate and detailed reconstructions. Our findings reveal that VSR models can perform remarkably well even on sequences that lack precise spatial alignment. Given this observation, we propose a simple yet practical approach to align LR images without involving fine-tuning or generating 'smooth' trajectory from the trained 3D models over LR images. The experimental results show that the surprisingly simple algorithms can achieve the state-of-the-art results of 3D super-resolution tasks on standard benchmark datasets, such as the NeRF-synthetic and MipNeRF-360 datasets. Project page: https://ko-lani.github.io/Sequence-Matters
title Sequence Matters: Harnessing Video Models in 3D Super-Resolution
topic Computer Vision and Pattern Recognition
68U10, 68T10
I.4.5; I.2.10
url https://arxiv.org/abs/2412.11525