Robustifying Diffusion-Denoised Smoothing Against Covariate Shift

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hedayatnia, Ali, Tavassolipour, Mostafa, Araabi, Babak Nadjar, Vahabie, Abdol-Hossein
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914036207058944
author Hedayatnia, Ali
Tavassolipour, Mostafa
Araabi, Babak Nadjar
Vahabie, Abdol-Hossein
author_facet Hedayatnia, Ali
Tavassolipour, Mostafa
Araabi, Babak Nadjar
Vahabie, Abdol-Hossein
contents Randomized smoothing is a well-established method for achieving certified robustness against l2-adversarial perturbations. By incorporating a denoiser before the base classifier, pretrained classifiers can be seamlessly integrated into randomized smoothing without significant performance degradation. Among existing methods, Diffusion Denoised Smoothing - where a pretrained denoising diffusion model serves as the denoiser - has produced state-of-the-art results. However, we show that employing a denoising diffusion model introduces a covariate shift via misestimation of the added noise, ultimately degrading the smoothed classifier's performance. To address this issue, we propose a novel adversarial objective function focused on the added noise of the denoising diffusion model. This approach is inspired by our understanding of the origin of the covariate shift. Our goal is to train the base classifier to ensure it is robust against the covariate shift introduced by the denoiser. Our method significantly improves certified accuracy across three standard classification benchmarks - MNIST, CIFAR-10, and ImageNet - achieving new state-of-the-art performance in l2-adversarial perturbations. Our implementation is publicly available at https://github.com/ahedayat/Robustifying-DDS-Against-Covariate-Shift
format Preprint
id arxiv_https___arxiv_org_abs_2509_10913
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Robustifying Diffusion-Denoised Smoothing Against Covariate Shift
Hedayatnia, Ali
Tavassolipour, Mostafa
Araabi, Babak Nadjar
Vahabie, Abdol-Hossein
Machine Learning
Computer Vision and Pattern Recognition
Randomized smoothing is a well-established method for achieving certified robustness against l2-adversarial perturbations. By incorporating a denoiser before the base classifier, pretrained classifiers can be seamlessly integrated into randomized smoothing without significant performance degradation. Among existing methods, Diffusion Denoised Smoothing - where a pretrained denoising diffusion model serves as the denoiser - has produced state-of-the-art results. However, we show that employing a denoising diffusion model introduces a covariate shift via misestimation of the added noise, ultimately degrading the smoothed classifier's performance. To address this issue, we propose a novel adversarial objective function focused on the added noise of the denoising diffusion model. This approach is inspired by our understanding of the origin of the covariate shift. Our goal is to train the base classifier to ensure it is robust against the covariate shift introduced by the denoiser. Our method significantly improves certified accuracy across three standard classification benchmarks - MNIST, CIFAR-10, and ImageNet - achieving new state-of-the-art performance in l2-adversarial perturbations. Our implementation is publicly available at https://github.com/ahedayat/Robustifying-DDS-Against-Covariate-Shift
title Robustifying Diffusion-Denoised Smoothing Against Covariate Shift
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.10913