PIV3CAMS: a multi-camera dataset for multiple computer vision problems and its application to novel view-point synthesis

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Kim, Sohyeong, Danelljan, Martin, Timofte, Radu, Van Gool, Luc, Thiran, Jean-Philippe
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911969044332544
author Kim, Sohyeong
Danelljan, Martin
Timofte, Radu
Van Gool, Luc
Thiran, Jean-Philippe
author_facet Kim, Sohyeong
Danelljan, Martin
Timofte, Radu
Van Gool, Luc
Thiran, Jean-Philippe
contents The modern approaches for computer vision tasks significantly rely on machine learning, which requires a large number of quality images. While there is a plethora of image datasets with a single type of images, there is a lack of datasets collected from multiple cameras. In this thesis, we introduce Paired Image and Video data from three CAMeraS, namely PIV3CAMS, aimed at multiple computer vision tasks. The PIV3CAMS dataset consists of 8385 pairs of images and 82 pairs of videos taken from three different cameras: Canon D5 Mark IV, Huawei P20, and ZED stereo camera. The dataset includes various indoor and outdoor scenes from different locations in Zurich (Switzerland) and Cheonan (South Korea). Some of the computer vision applications that can benefit from the PIV3CAMS dataset are image/video enhancement, view interpolation, image matching, and much more. We provide a careful explanation of the data collection process and detailed analysis of the data. The second part of this thesis studies the usage of depth information in the view synthesizing task. In addition to the regeneration of a current state-of-the-art algorithm, we investigate several proposed alternative models that integrate depth information geometrically. Through extensive experiments, we show that the effect of depth is crucial in small view changes. Finally, we apply our model to the introduced PIV3CAMS dataset to synthesize novel target views as an example application of PIV3CAMS.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18695
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PIV3CAMS: a multi-camera dataset for multiple computer vision problems and its application to novel view-point synthesis
Kim, Sohyeong
Danelljan, Martin
Timofte, Radu
Van Gool, Luc
Thiran, Jean-Philippe
Computer Vision and Pattern Recognition
The modern approaches for computer vision tasks significantly rely on machine learning, which requires a large number of quality images. While there is a plethora of image datasets with a single type of images, there is a lack of datasets collected from multiple cameras. In this thesis, we introduce Paired Image and Video data from three CAMeraS, namely PIV3CAMS, aimed at multiple computer vision tasks. The PIV3CAMS dataset consists of 8385 pairs of images and 82 pairs of videos taken from three different cameras: Canon D5 Mark IV, Huawei P20, and ZED stereo camera. The dataset includes various indoor and outdoor scenes from different locations in Zurich (Switzerland) and Cheonan (South Korea). Some of the computer vision applications that can benefit from the PIV3CAMS dataset are image/video enhancement, view interpolation, image matching, and much more. We provide a careful explanation of the data collection process and detailed analysis of the data. The second part of this thesis studies the usage of depth information in the view synthesizing task. In addition to the regeneration of a current state-of-the-art algorithm, we investigate several proposed alternative models that integrate depth information geometrically. Through extensive experiments, we show that the effect of depth is crucial in small view changes. Finally, we apply our model to the introduced PIV3CAMS dataset to synthesize novel target views as an example application of PIV3CAMS.
title PIV3CAMS: a multi-camera dataset for multiple computer vision problems and its application to novel view-point synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18695