Resolution Invariant Autoencoder

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Patel, Ashay, Antonelli, Michela, Ourselin, Sebastien, Cardoso, M. Jorge
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929757316186112
author Patel, Ashay
Antonelli, Michela
Ourselin, Sebastien
Cardoso, M. Jorge
author_facet Patel, Ashay
Antonelli, Michela
Ourselin, Sebastien
Cardoso, M. Jorge
contents Deep learning has significantly advanced medical imaging analysis, yet variations in image resolution remain an overlooked challenge. Most methods address this by resampling images, leading to either information loss or computational inefficiencies. While solutions exist for specific tasks, no unified approach has been proposed. We introduce a resolution-invariant autoencoder that adapts spatial resizing at each layer in the network via a learned variable resizing process, replacing fixed spatial down/upsampling at the traditional factor of 2. This ensures a consistent latent space resolution, regardless of input or output resolution. Our model enables various downstream tasks to be performed on an image latent whilst maintaining performance across different resolutions, overcoming the shortfalls of traditional methods. We demonstrate its effectiveness in uncertainty-aware super-resolution, classification, and generative modelling tasks and show how our method outperforms conventional baselines with minimal performance loss across resolutions.
format Preprint
id arxiv_https___arxiv_org_abs_2503_09828
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Resolution Invariant Autoencoder
Patel, Ashay
Antonelli, Michela
Ourselin, Sebastien
Cardoso, M. Jorge
Computer Vision and Pattern Recognition
Image and Video Processing
Deep learning has significantly advanced medical imaging analysis, yet variations in image resolution remain an overlooked challenge. Most methods address this by resampling images, leading to either information loss or computational inefficiencies. While solutions exist for specific tasks, no unified approach has been proposed. We introduce a resolution-invariant autoencoder that adapts spatial resizing at each layer in the network via a learned variable resizing process, replacing fixed spatial down/upsampling at the traditional factor of 2. This ensures a consistent latent space resolution, regardless of input or output resolution. Our model enables various downstream tasks to be performed on an image latent whilst maintaining performance across different resolutions, overcoming the shortfalls of traditional methods. We demonstrate its effectiveness in uncertainty-aware super-resolution, classification, and generative modelling tasks and show how our method outperforms conventional baselines with minimal performance loss across resolutions.
title Resolution Invariant Autoencoder
topic Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2503.09828