Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dalmonte, Francesco, Bayar, Emirhan, Akbas, Emre, Georgescu, Mariana-Iuliana
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912499880689664
author Dalmonte, Francesco
Bayar, Emirhan
Akbas, Emre
Georgescu, Mariana-Iuliana
author_facet Dalmonte, Francesco
Bayar, Emirhan
Akbas, Emre
Georgescu, Mariana-Iuliana
contents Anomaly detection in medical images is an important yet challenging task due to the diversity of possible anomalies and the practical impossibility of collecting comprehensively annotated data sets. In this work, we tackle unsupervised medical anomaly detection proposing a modernized autoencoder-based framework, the Q-Former Autoencoder, that leverages state-of-the-art pretrained vision foundation models, such as DINO, DINOv2 and Masked Autoencoder. Instead of training encoders from scratch, we directly utilize frozen vision foundation models as feature extractors, enabling rich, multi-stage, high-level representations without domain-specific fine-tuning. We propose the usage of the Q-Former architecture as the bottleneck, which enables the control of the length of the reconstruction sequence, while efficiently aggregating multiscale features. Additionally, we incorporate a perceptual loss computed using features from a pretrained Masked Autoencoder, guiding the reconstruction towards semantically meaningful structures. Our framework is evaluated on four diverse medical anomaly detection benchmarks, achieving state-of-the-art results on BraTS2021, RESC, and RSNA. Our results highlight the potential of vision foundation model encoders, pretrained on natural images, to generalize effectively to medical image analysis tasks without further fine-tuning. We release the code and models at https://github.com/emirhanbayar/QFAE.
format Preprint
id arxiv_https___arxiv_org_abs_2507_18481
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
Dalmonte, Francesco
Bayar, Emirhan
Akbas, Emre
Georgescu, Mariana-Iuliana
Computer Vision and Pattern Recognition
Anomaly detection in medical images is an important yet challenging task due to the diversity of possible anomalies and the practical impossibility of collecting comprehensively annotated data sets. In this work, we tackle unsupervised medical anomaly detection proposing a modernized autoencoder-based framework, the Q-Former Autoencoder, that leverages state-of-the-art pretrained vision foundation models, such as DINO, DINOv2 and Masked Autoencoder. Instead of training encoders from scratch, we directly utilize frozen vision foundation models as feature extractors, enabling rich, multi-stage, high-level representations without domain-specific fine-tuning. We propose the usage of the Q-Former architecture as the bottleneck, which enables the control of the length of the reconstruction sequence, while efficiently aggregating multiscale features. Additionally, we incorporate a perceptual loss computed using features from a pretrained Masked Autoencoder, guiding the reconstruction towards semantically meaningful structures. Our framework is evaluated on four diverse medical anomaly detection benchmarks, achieving state-of-the-art results on BraTS2021, RESC, and RSNA. Our results highlight the potential of vision foundation model encoders, pretrained on natural images, to generalize effectively to medical image analysis tasks without further fine-tuning. We release the code and models at https://github.com/emirhanbayar/QFAE.
title Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.18481