Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jin, Gaojie, Huang, Tianjin, Mu, Ronghui, Huang, Xiaowei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913750318055424
author Jin, Gaojie
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
author_facet Jin, Gaojie
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
contents Recent studies have identified a critical challenge in deep neural networks (DNNs) known as ``robust fairness", where models exhibit significant disparities in robust accuracy across different classes. While prior work has attempted to address this issue in adversarial robustness, the study of worst-class certified robustness for smoothed classifiers remains unexplored. Our work bridges this gap by developing a PAC-Bayesian bound for the worst-class error of smoothed classifiers. Through theoretical analysis, we demonstrate that the largest eigenvalue of the smoothed confusion matrix fundamentally influences the worst-class error of smoothed classifiers. Based on this insight, we introduce a regularization method that optimizes the largest eigenvalue of smoothed confusion matrix to enhance worst-class accuracy of the smoothed classifier and further improve its worst-class certified robustness. We provide extensive experimental validation across multiple datasets and model architectures to demonstrate the effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers
Jin, Gaojie
Huang, Tianjin
Mu, Ronghui
Huang, Xiaowei
Machine Learning
Recent studies have identified a critical challenge in deep neural networks (DNNs) known as ``robust fairness", where models exhibit significant disparities in robust accuracy across different classes. While prior work has attempted to address this issue in adversarial robustness, the study of worst-class certified robustness for smoothed classifiers remains unexplored. Our work bridges this gap by developing a PAC-Bayesian bound for the worst-class error of smoothed classifiers. Through theoretical analysis, we demonstrate that the largest eigenvalue of the smoothed confusion matrix fundamentally influences the worst-class error of smoothed classifiers. Based on this insight, we introduce a regularization method that optimizes the largest eigenvalue of smoothed confusion matrix to enhance worst-class accuracy of the smoothed classifier and further improve its worst-class certified robustness. We provide extensive experimental validation across multiple datasets and model architectures to demonstrate the effectiveness of our approach.
title Principal Eigenvalue Regularization for Improved Worst-Class Certified Robustness of Smoothed Classifiers
topic Machine Learning
url https://arxiv.org/abs/2503.17172