Selective Learning: Towards Robust Calibration with Dynamic Regularization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Han, Zongbo, Yang, Yifeng, Zhang, Changqing, Zhang, Linjun, Zhou, Joey Tianyi, Hu, Qinghua
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909252875976704
author Han, Zongbo
Yang, Yifeng
Zhang, Changqing
Zhang, Linjun
Zhou, Joey Tianyi
Hu, Qinghua
author_facet Han, Zongbo
Yang, Yifeng
Zhang, Changqing
Zhang, Linjun
Zhou, Joey Tianyi
Hu, Qinghua
contents Miscalibration in deep learning refers to there is a discrepancy between the predicted confidence and performance. This problem usually arises due to the overfitting problem, which is characterized by learning everything presented in the training set, resulting in overconfident predictions during testing. Existing methods typically address overfitting and mitigate the miscalibration by adding a maximum-entropy regularizer to the objective function. The objective can be understood as seeking a model that fits the ground-truth labels by increasing the confidence while also maximizing the entropy of predicted probabilities by decreasing the confidence. However, previous methods lack clear guidance on confidence adjustment, leading to conflicting objectives (increasing but also decreasing confidence). Therefore, we introduce a method called Dynamic Regularization (DReg), which aims to learn what should be learned during training thereby circumventing the confidence adjusting trade-off. At a high level, DReg aims to obtain a more reliable model capable of acknowledging what it knows and does not know. Specifically, DReg effectively fits the labels for in-distribution samples (samples that should be learned) while applying regularization dynamically to samples beyond model capabilities (e.g., outliers), thereby obtaining a robust calibrated model especially on the samples beyond model capabilities. Both theoretical and empirical analyses sufficiently demonstrate the superiority of DReg compared with previous methods.
format Preprint
id arxiv_https___arxiv_org_abs_2402_08384
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Selective Learning: Towards Robust Calibration with Dynamic Regularization
Han, Zongbo
Yang, Yifeng
Zhang, Changqing
Zhang, Linjun
Zhou, Joey Tianyi
Hu, Qinghua
Machine Learning
Artificial Intelligence
Miscalibration in deep learning refers to there is a discrepancy between the predicted confidence and performance. This problem usually arises due to the overfitting problem, which is characterized by learning everything presented in the training set, resulting in overconfident predictions during testing. Existing methods typically address overfitting and mitigate the miscalibration by adding a maximum-entropy regularizer to the objective function. The objective can be understood as seeking a model that fits the ground-truth labels by increasing the confidence while also maximizing the entropy of predicted probabilities by decreasing the confidence. However, previous methods lack clear guidance on confidence adjustment, leading to conflicting objectives (increasing but also decreasing confidence). Therefore, we introduce a method called Dynamic Regularization (DReg), which aims to learn what should be learned during training thereby circumventing the confidence adjusting trade-off. At a high level, DReg aims to obtain a more reliable model capable of acknowledging what it knows and does not know. Specifically, DReg effectively fits the labels for in-distribution samples (samples that should be learned) while applying regularization dynamically to samples beyond model capabilities (e.g., outliers), thereby obtaining a robust calibrated model especially on the samples beyond model capabilities. Both theoretical and empirical analyses sufficiently demonstrate the superiority of DReg compared with previous methods.
title Selective Learning: Towards Robust Calibration with Dynamic Regularization
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2402.08384