Saved in:
Bibliographic Details
Main Author: Du, Zhehang
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2311.02622
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917811403620352
author Du, Zhehang
author_facet Du, Zhehang
contents Neural networks often exhibit simplicity bias, favoring simpler features over more complex ones, even when both are equally predictive. We introduce a novel method called imbalanced label coupling to explore and extend this simplicity bias across multiple hierarchical levels. Our approach demonstrates that trained networks sequentially consider features of increasing complexity based on their correlation with labels in the training set, regardless of their actual predictive power. For example, in CIFAR-10, simple spurious features can cause misclassifications where most cats are predicted as dogs and most trucks as automobiles. We empirically show that last-layer retraining with target data distribution \citep{kirichenko2022last} is insufficient to fully recover core features when spurious features perfectly correlate with target labels in our synthetic datasets. Our findings deepen the understanding of the implicit biases inherent in neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2311_02622
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Hierarchical Simplicity Bias of Neural Networks
Du, Zhehang
Machine Learning
Computer Vision and Pattern Recognition
Neural networks often exhibit simplicity bias, favoring simpler features over more complex ones, even when both are equally predictive. We introduce a novel method called imbalanced label coupling to explore and extend this simplicity bias across multiple hierarchical levels. Our approach demonstrates that trained networks sequentially consider features of increasing complexity based on their correlation with labels in the training set, regardless of their actual predictive power. For example, in CIFAR-10, simple spurious features can cause misclassifications where most cats are predicted as dogs and most trucks as automobiles. We empirically show that last-layer retraining with target data distribution \citep{kirichenko2022last} is insufficient to fully recover core features when spurious features perfectly correlate with target labels in our synthetic datasets. Our findings deepen the understanding of the implicit biases inherent in neural networks.
title Hierarchical Simplicity Bias of Neural Networks
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.02622