Towards Certification of Uncertainty Calibration under Adversarial Attacks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Emde, Cornelius, Pinto, Francesco, Lukasiewicz, Thomas, Torr, Philip H. S., Bibi, Adel
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929730552332288
author Emde, Cornelius
Pinto, Francesco
Lukasiewicz, Thomas
Torr, Philip H. S.
Bibi, Adel
author_facet Emde, Cornelius
Pinto, Francesco
Lukasiewicz, Thomas
Torr, Philip H. S.
Bibi, Adel
contents Since neural classifiers are known to be sensitive to adversarial perturbations that alter their accuracy, \textit{certification methods} have been developed to provide provable guarantees on the insensitivity of their predictions to such perturbations. Furthermore, in safety-critical applications, the frequentist interpretation of the confidence of a classifier (also known as model calibration) can be of utmost importance. This property can be measured via the Brier score or the expected calibration error. We show that attacks can significantly harm calibration, and thus propose certified calibration as worst-case bounds on calibration under adversarial perturbations. Specifically, we produce analytic bounds for the Brier score and approximate bounds via the solution of a mixed-integer program on the expected calibration error. Finally, we propose novel calibration attacks and demonstrate how they can improve model calibration through \textit{adversarial calibration training}.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13922
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Towards Certification of Uncertainty Calibration under Adversarial Attacks
Emde, Cornelius
Pinto, Francesco
Lukasiewicz, Thomas
Torr, Philip H. S.
Bibi, Adel
Machine Learning
Since neural classifiers are known to be sensitive to adversarial perturbations that alter their accuracy, \textit{certification methods} have been developed to provide provable guarantees on the insensitivity of their predictions to such perturbations. Furthermore, in safety-critical applications, the frequentist interpretation of the confidence of a classifier (also known as model calibration) can be of utmost importance. This property can be measured via the Brier score or the expected calibration error. We show that attacks can significantly harm calibration, and thus propose certified calibration as worst-case bounds on calibration under adversarial perturbations. Specifically, we produce analytic bounds for the Brier score and approximate bounds via the solution of a mixed-integer program on the expected calibration error. Finally, we propose novel calibration attacks and demonstrate how they can improve model calibration through \textit{adversarial calibration training}.
title Towards Certification of Uncertainty Calibration under Adversarial Attacks
topic Machine Learning
url https://arxiv.org/abs/2405.13922