Measures of Information Reflect Memorization Patterns

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bansal, Rachit, Pruthi, Danish, Belinkov, Yonatan
Natura: Preprint
Pubblicazione: 2022
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916112336158720
author Bansal, Rachit
Pruthi, Danish
Belinkov, Yonatan
author_facet Bansal, Rachit
Pruthi, Danish
Belinkov, Yonatan
contents Neural networks are known to exploit spurious artifacts (or shortcuts) that co-occur with a target label, exhibiting heuristic memorization. On the other hand, networks have been shown to memorize training examples, resulting in example-level memorization. These kinds of memorization impede generalization of networks beyond their training distributions. Detecting such memorization could be challenging, often requiring researchers to curate tailored test sets. In this work, we hypothesize -- and subsequently show -- that the diversity in the activation patterns of different neurons is reflective of model generalization and memorization. We quantify the diversity in the neural activations through information-theoretic measures and find support for our hypothesis on experiments spanning several natural language and vision tasks. Importantly, we discover that information organization points to the two forms of memorization, even for neural activations computed on unlabelled in-distribution examples. Lastly, we demonstrate the utility of our findings for the problem of model selection. The associated code and other resources for this work are available at https://rachitbansal.github.io/information-measures.
format Preprint
id arxiv_https___arxiv_org_abs_2210_09404
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Measures of Information Reflect Memorization Patterns
Bansal, Rachit
Pruthi, Danish
Belinkov, Yonatan
Machine Learning
Information Theory
Neural networks are known to exploit spurious artifacts (or shortcuts) that co-occur with a target label, exhibiting heuristic memorization. On the other hand, networks have been shown to memorize training examples, resulting in example-level memorization. These kinds of memorization impede generalization of networks beyond their training distributions. Detecting such memorization could be challenging, often requiring researchers to curate tailored test sets. In this work, we hypothesize -- and subsequently show -- that the diversity in the activation patterns of different neurons is reflective of model generalization and memorization. We quantify the diversity in the neural activations through information-theoretic measures and find support for our hypothesis on experiments spanning several natural language and vision tasks. Importantly, we discover that information organization points to the two forms of memorization, even for neural activations computed on unlabelled in-distribution examples. Lastly, we demonstrate the utility of our findings for the problem of model selection. The associated code and other resources for this work are available at https://rachitbansal.github.io/information-measures.
title Measures of Information Reflect Memorization Patterns
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2210.09404