Decoupled Kullback-Leibler Divergence Loss

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cui, Jiequan, Tian, Zhuotao, Zhong, Zhisheng, Qi, Xiaojuan, Yu, Bei, Zhang, Hanwang
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929561458966528
author Cui, Jiequan
Tian, Zhuotao
Zhong, Zhisheng
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
author_facet Cui, Jiequan
Tian, Zhuotao
Zhong, Zhisheng
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
contents In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of 1) a weighted Mean Square Error (wMSE) loss and 2) a Cross-Entropy loss incorporating soft labels. Thanks to the decomposed formulation of DKL loss, we have identified two areas for improvement. Firstly, we address the limitation of KL/DKL in scenarios like knowledge distillation by breaking its asymmetric optimization property. This modification ensures that the $\mathbf{w}$MSE component is always effective during training, providing extra constructive cues. Secondly, we introduce class-wise global information into KL/DKL to mitigate bias from individual samples. With these two enhancements, we derive the Improved Kullback-Leibler (IKL) Divergence loss and evaluate its effectiveness by conducting experiments on CIFAR-10/100 and ImageNet datasets, focusing on adversarial training, and knowledge distillation tasks. The proposed approach achieves new state-of-the-art adversarial robustness on the public leaderboard -- RobustBench and competitive performance on knowledge distillation, demonstrating the substantial practical merits. Our code is available at https://github.com/jiequancui/DKL.
format Preprint
id arxiv_https___arxiv_org_abs_2305_13948
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Decoupled Kullback-Leibler Divergence Loss
Cui, Jiequan
Tian, Zhuotao
Zhong, Zhisheng
Qi, Xiaojuan
Yu, Bei
Zhang, Hanwang
Computer Vision and Pattern Recognition
Machine Learning
In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of 1) a weighted Mean Square Error (wMSE) loss and 2) a Cross-Entropy loss incorporating soft labels. Thanks to the decomposed formulation of DKL loss, we have identified two areas for improvement. Firstly, we address the limitation of KL/DKL in scenarios like knowledge distillation by breaking its asymmetric optimization property. This modification ensures that the $\mathbf{w}$MSE component is always effective during training, providing extra constructive cues. Secondly, we introduce class-wise global information into KL/DKL to mitigate bias from individual samples. With these two enhancements, we derive the Improved Kullback-Leibler (IKL) Divergence loss and evaluate its effectiveness by conducting experiments on CIFAR-10/100 and ImageNet datasets, focusing on adversarial training, and knowledge distillation tasks. The proposed approach achieves new state-of-the-art adversarial robustness on the public leaderboard -- RobustBench and competitive performance on knowledge distillation, demonstrating the substantial practical merits. Our code is available at https://github.com/jiequancui/DKL.
title Decoupled Kullback-Leibler Divergence Loss
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2305.13948