PIP: Positional-encoding Image Prior

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shabtay, Nimrod, Schwartz, Eli, Giryes, Raja
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916144025174016
author Shabtay, Nimrod
Schwartz, Eli
Giryes, Raja
author_facet Shabtay, Nimrod
Schwartz, Eli
Giryes, Raja
contents In Deep Image Prior (DIP), a Convolutional Neural Network (CNN) is fitted to map a latent space to a degraded (e.g. noisy) image but in the process learns to reconstruct the clean image. This phenomenon is attributed to CNN's internal image-prior. We revisit the DIP framework, examining it from the perspective of a neural implicit representation. Motivated by this perspective, we replace the random or learned latent with Fourier-Features (Positional Encoding). We show that thanks to the Fourier features properties, we can replace the convolution layers with simple pixel-level MLPs. We name this scheme ``Positional Encoding Image Prior" (PIP) and exhibit that it performs very similarly to DIP on various image-reconstruction tasks with much less parameters required. Additionally, we demonstrate that PIP can be easily extended to videos, where 3D-DIP struggles and suffers from instability. Code and additional examples for all tasks, including videos, are available on the project page https://nimrodshabtay.github.io/PIP/
format Preprint
id arxiv_https___arxiv_org_abs_2211_14298
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle PIP: Positional-encoding Image Prior
Shabtay, Nimrod
Schwartz, Eli
Giryes, Raja
Computer Vision and Pattern Recognition
Artificial Intelligence
In Deep Image Prior (DIP), a Convolutional Neural Network (CNN) is fitted to map a latent space to a degraded (e.g. noisy) image but in the process learns to reconstruct the clean image. This phenomenon is attributed to CNN's internal image-prior. We revisit the DIP framework, examining it from the perspective of a neural implicit representation. Motivated by this perspective, we replace the random or learned latent with Fourier-Features (Positional Encoding). We show that thanks to the Fourier features properties, we can replace the convolution layers with simple pixel-level MLPs. We name this scheme ``Positional Encoding Image Prior" (PIP) and exhibit that it performs very similarly to DIP on various image-reconstruction tasks with much less parameters required. Additionally, we demonstrate that PIP can be easily extended to videos, where 3D-DIP struggles and suffers from instability. Code and additional examples for all tasks, including videos, are available on the project page https://nimrodshabtay.github.io/PIP/
title PIP: Positional-encoding Image Prior
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2211.14298