Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Oh, Yeongtak, Lee, Jonghyun, Choi, Jooyoung, Jung, Dahuin, Hwang, Uiwon, Yoon, Sungroh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911951774285824
author Oh, Yeongtak
Lee, Jonghyun
Choi, Jooyoung
Jung, Dahuin
Hwang, Uiwon
Yoon, Sungroh
author_facet Oh, Yeongtak
Lee, Jonghyun
Choi, Jooyoung
Jung, Dahuin
Hwang, Uiwon
Yoon, Sungroh
contents Test-time adaptation (TTA) addresses the unforeseen distribution shifts occurring during test time. In TTA, performance, memory consumption, and time consumption are crucial considerations. A recent diffusion-based TTA approach for restoring corrupted images involves image-level updates. However, using pixel space diffusion significantly increases resource requirements compared to conventional model updating TTA approaches, revealing limitations as a TTA method. To address this, we propose a novel TTA method that leverages an image editing model based on a latent diffusion model (LDM) and fine-tunes it using our newly introduced corruption modeling scheme. This scheme enhances the robustness of the diffusion model against distribution shifts by creating (clean, corrupted) image pairs and fine-tuning the model to edit corrupted images into clean ones. Moreover, we introduce a distilled variant to accelerate the model for corruption editing using only 4 network function evaluations (NFEs). We extensively validated our method across various architectures and datasets including image and video domains. Our model achieves the best performance with a 100 times faster runtime than that of a diffusion-based baseline. Furthermore, it is three times faster than the previous model updating TTA method that utilizes data augmentation, making an image-level updating approach more feasible.
format Preprint
id arxiv_https___arxiv_org_abs_2403_10911
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
Oh, Yeongtak
Lee, Jonghyun
Choi, Jooyoung
Jung, Dahuin
Hwang, Uiwon
Yoon, Sungroh
Computer Vision and Pattern Recognition
Test-time adaptation (TTA) addresses the unforeseen distribution shifts occurring during test time. In TTA, performance, memory consumption, and time consumption are crucial considerations. A recent diffusion-based TTA approach for restoring corrupted images involves image-level updates. However, using pixel space diffusion significantly increases resource requirements compared to conventional model updating TTA approaches, revealing limitations as a TTA method. To address this, we propose a novel TTA method that leverages an image editing model based on a latent diffusion model (LDM) and fine-tunes it using our newly introduced corruption modeling scheme. This scheme enhances the robustness of the diffusion model against distribution shifts by creating (clean, corrupted) image pairs and fine-tuning the model to edit corrupted images into clean ones. Moreover, we introduce a distilled variant to accelerate the model for corruption editing using only 4 network function evaluations (NFEs). We extensively validated our method across various architectures and datasets including image and video domains. Our model achieves the best performance with a 100 times faster runtime than that of a diffusion-based baseline. Furthermore, it is three times faster than the previous model updating TTA method that utilizes data augmentation, making an image-level updating approach more feasible.
title Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2403.10911