Rethinking the Harmonic Loss via Non-Euclidean Distance Layers

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Miller-Golub, Maxwell, Coil, Collin, Faber, Kamil, Pietron, Marcin, Zheng, Panpan, Minervini, Pasquale, Corizzo, Roberto
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908999311425536
author Miller-Golub, Maxwell
Coil, Collin
Faber, Kamil
Pietron, Marcin
Zheng, Panpan
Minervini, Pasquale
Corizzo, Roberto
author_facet Miller-Golub, Maxwell
Coil, Collin
Faber, Kamil
Pietron, Marcin
Zheng, Panpan
Minervini, Pasquale
Corizzo, Roberto
contents Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbounded weight growth, and inefficiencies that can contribute to costly training dynamics. The harmonic loss is a distance-based alternative grounded in Euclidean geometry that improves interpretability and mitigates phenomena such as grokking, or delayed generalization on the test set. However, the study of harmonic loss remains narrow: only Euclidean distance is explored, and no systematic evaluation of computational efficiency or sustainability was conducted. We extend harmonic loss by systematically investigating a broad spectrum of distance metrics as replacements for the Euclidean distance. We comprehensively evaluate distance-tailored harmonic losses on both vision backbones and large language models. Our analysis is framed around a three-way evaluation of model performance, interpretability, and sustainability. On vision tasks, cosine distances provide the most favorable trade-off, consistently improving accuracy while lowering carbon emissions, whereas Bray-Curtis and Mahalanobis further enhance interpretability at varying efficiency costs. On language models, cosine-based harmonic losses improve gradient and learning stability, strengthen representation structure, and reduce emissions relative to cross-entropy and Euclidean heads. Our code is available at: https://anonymous.4open.science/r/rethinking-harmonic-loss-5BAB/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_10225
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking the Harmonic Loss via Non-Euclidean Distance Layers
Miller-Golub, Maxwell
Coil, Collin
Faber, Kamil
Pietron, Marcin
Zheng, Panpan
Minervini, Pasquale
Corizzo, Roberto
Machine Learning
Artificial Intelligence
Cross-entropy loss has long been the standard choice for training deep neural networks, yet it suffers from interpretability limitations, unbounded weight growth, and inefficiencies that can contribute to costly training dynamics. The harmonic loss is a distance-based alternative grounded in Euclidean geometry that improves interpretability and mitigates phenomena such as grokking, or delayed generalization on the test set. However, the study of harmonic loss remains narrow: only Euclidean distance is explored, and no systematic evaluation of computational efficiency or sustainability was conducted. We extend harmonic loss by systematically investigating a broad spectrum of distance metrics as replacements for the Euclidean distance. We comprehensively evaluate distance-tailored harmonic losses on both vision backbones and large language models. Our analysis is framed around a three-way evaluation of model performance, interpretability, and sustainability. On vision tasks, cosine distances provide the most favorable trade-off, consistently improving accuracy while lowering carbon emissions, whereas Bray-Curtis and Mahalanobis further enhance interpretability at varying efficiency costs. On language models, cosine-based harmonic losses improve gradient and learning stability, strengthen representation structure, and reduce emissions relative to cross-entropy and Euclidean heads. Our code is available at: https://anonymous.4open.science/r/rethinking-harmonic-loss-5BAB/.
title Rethinking the Harmonic Loss via Non-Euclidean Distance Layers
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2603.10225