Deep Companion Learning: Enhancing Generalization Through Historical Consistency

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Ruizhao, Saligrama, Venkatesh
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914888813641728
author Zhu, Ruizhao
Saligrama, Venkatesh
author_facet Zhu, Ruizhao
Saligrama, Venkatesh
contents We propose Deep Companion Learning (DCL), a novel training method for Deep Neural Networks (DNNs) that enhances generalization by penalizing inconsistent model predictions compared to its historical performance. To achieve this, we train a deep-companion model (DCM), by using previous versions of the model to provide forecasts on new inputs. This companion model deciphers a meaningful latent semantic structure within the data, thereby providing targeted supervision that encourages the primary model to address the scenarios it finds most challenging. We validate our approach through both theoretical analysis and extensive experimentation, including ablation studies, on a variety of benchmark datasets (CIFAR-100, Tiny-ImageNet, ImageNet-1K) using diverse architectural models (ShuffleNetV2, ResNet, Vision Transformer, etc.), demonstrating state-of-the-art performance.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18821
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Deep Companion Learning: Enhancing Generalization Through Historical Consistency
Zhu, Ruizhao
Saligrama, Venkatesh
Computer Vision and Pattern Recognition
Machine Learning
We propose Deep Companion Learning (DCL), a novel training method for Deep Neural Networks (DNNs) that enhances generalization by penalizing inconsistent model predictions compared to its historical performance. To achieve this, we train a deep-companion model (DCM), by using previous versions of the model to provide forecasts on new inputs. This companion model deciphers a meaningful latent semantic structure within the data, thereby providing targeted supervision that encourages the primary model to address the scenarios it finds most challenging. We validate our approach through both theoretical analysis and extensive experimentation, including ablation studies, on a variety of benchmark datasets (CIFAR-100, Tiny-ImageNet, ImageNet-1K) using diverse architectural models (ShuffleNetV2, ResNet, Vision Transformer, etc.), demonstrating state-of-the-art performance.
title Deep Companion Learning: Enhancing Generalization Through Historical Consistency
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2407.18821