On the Use of Anchoring for Training Vision Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Narayanaswamy, Vivek, Thopalli, Kowshik, Anirudh, Rushil, Mubarka, Yamen, Sakla, Wesam, Thiagarajan, Jayaraman J.
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909214721441792
author Narayanaswamy, Vivek
Thopalli, Kowshik
Anirudh, Rushil
Mubarka, Yamen
Sakla, Wesam
Thiagarajan, Jayaraman J.
author_facet Narayanaswamy, Vivek
Thopalli, Kowshik
Anirudh, Rushil
Mubarka, Yamen
Sakla, Wesam
Thiagarajan, Jayaraman J.
contents Anchoring is a recent, architecture-agnostic principle for training deep neural networks that has been shown to significantly improve uncertainty estimation, calibration, and extrapolation capabilities. In this paper, we systematically explore anchoring as a general protocol for training vision models, providing fundamental insights into its training and inference processes and their implications for generalization and safety. Despite its promise, we identify a critical problem in anchored training that can lead to an increased risk of learning undesirable shortcuts, thereby limiting its generalization capabilities. To address this, we introduce a new anchored training protocol that employs a simple regularizer to mitigate this issue and significantly enhances generalization. We empirically evaluate our proposed approach across datasets and architectures of varying scales and complexities, demonstrating substantial performance gains in generalization and safety metrics compared to the standard training protocol.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00529
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Use of Anchoring for Training Vision Models
Narayanaswamy, Vivek
Thopalli, Kowshik
Anirudh, Rushil
Mubarka, Yamen
Sakla, Wesam
Thiagarajan, Jayaraman J.
Machine Learning
Computer Vision and Pattern Recognition
Anchoring is a recent, architecture-agnostic principle for training deep neural networks that has been shown to significantly improve uncertainty estimation, calibration, and extrapolation capabilities. In this paper, we systematically explore anchoring as a general protocol for training vision models, providing fundamental insights into its training and inference processes and their implications for generalization and safety. Despite its promise, we identify a critical problem in anchored training that can lead to an increased risk of learning undesirable shortcuts, thereby limiting its generalization capabilities. To address this, we introduce a new anchored training protocol that employs a simple regularizer to mitigate this issue and significantly enhances generalization. We empirically evaluate our proposed approach across datasets and architectures of varying scales and complexities, demonstrating substantial performance gains in generalization and safety metrics compared to the standard training protocol.
title On the Use of Anchoring for Training Vision Models
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2406.00529