Minimax Generalized Cross-Entropy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bondugula, Kartheek, Mazuelas, Santiago, Pérez, Aritz, Liu, Anqi
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908996701519872
author Bondugula, Kartheek
Mazuelas, Santiago
Pérez, Aritz
Liu, Anqi
author_facet Bondugula, Kartheek
Mazuelas, Santiago
Pérez, Aritz
Liu, Anqi
contents Loss functions play a central role in supervised classification. Cross-entropy (CE) is widely used, whereas the mean absolute error (MAE) loss can offer robustness but is difficult to optimize. Interpolating between the CE and MAE losses, generalized cross-entropy (GCE) has recently been introduced to provide a trade-off between optimization difficulty and robustness. Existing formulations of GCE result in a non-convex optimization over classification margins that is prone to underfitting, leading to poor performances with complex datasets. In this paper, we propose a minimax formulation of generalized cross-entropy (MGCE) that results in a convex optimization over classification margins. Moreover, we show that MGCEs can provide an upper bound on the classification error. The proposed bilevel convex optimization can be efficiently implemented using stochastic gradient computed via implicit differentiation. Using benchmark datasets, we show that MGCE achieves strong accuracy, faster convergence, and better calibration, especially in the presence of label noise.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19874
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Minimax Generalized Cross-Entropy
Bondugula, Kartheek
Mazuelas, Santiago
Pérez, Aritz
Liu, Anqi
Machine Learning
Loss functions play a central role in supervised classification. Cross-entropy (CE) is widely used, whereas the mean absolute error (MAE) loss can offer robustness but is difficult to optimize. Interpolating between the CE and MAE losses, generalized cross-entropy (GCE) has recently been introduced to provide a trade-off between optimization difficulty and robustness. Existing formulations of GCE result in a non-convex optimization over classification margins that is prone to underfitting, leading to poor performances with complex datasets. In this paper, we propose a minimax formulation of generalized cross-entropy (MGCE) that results in a convex optimization over classification margins. Moreover, we show that MGCEs can provide an upper bound on the classification error. The proposed bilevel convex optimization can be efficiently implemented using stochastic gradient computed via implicit differentiation. Using benchmark datasets, we show that MGCE achieves strong accuracy, faster convergence, and better calibration, especially in the presence of label noise.
title Minimax Generalized Cross-Entropy
topic Machine Learning
url https://arxiv.org/abs/2603.19874