When Representations Align: Universality in Representation Learning Dynamics

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: van Rossem, Loek, Saxe, Andrew M.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916313040945152
author van Rossem, Loek
Saxe, Andrew M.
author_facet van Rossem, Loek
Saxe, Andrew M.
contents Deep neural networks come in many sizes and architectures. The choice of architecture, in conjunction with the dataset and learning algorithm, is commonly understood to affect the learned neural representations. Yet, recent results have shown that different architectures learn representations with striking qualitative similarities. Here we derive an effective theory of representation learning under the assumption that the encoding map from input to hidden representation and the decoding map from representation to output are arbitrary smooth functions. This theory schematizes representation learning dynamics in the regime of complex, large architectures, where hidden representations are not strongly constrained by the parametrization. We show through experiments that the effective theory describes aspects of representation learning dynamics across a range of deep networks with different activation functions and architectures, and exhibits phenomena similar to the "rich" and "lazy" regime. While many network behaviors depend quantitatively on architecture, our findings point to certain behaviors that are widely conserved once models are sufficiently flexible.
format Preprint
id arxiv_https___arxiv_org_abs_2402_09142
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle When Representations Align: Universality in Representation Learning Dynamics
van Rossem, Loek
Saxe, Andrew M.
Machine Learning
Neurons and Cognition
Deep neural networks come in many sizes and architectures. The choice of architecture, in conjunction with the dataset and learning algorithm, is commonly understood to affect the learned neural representations. Yet, recent results have shown that different architectures learn representations with striking qualitative similarities. Here we derive an effective theory of representation learning under the assumption that the encoding map from input to hidden representation and the decoding map from representation to output are arbitrary smooth functions. This theory schematizes representation learning dynamics in the regime of complex, large architectures, where hidden representations are not strongly constrained by the parametrization. We show through experiments that the effective theory describes aspects of representation learning dynamics across a range of deep networks with different activation functions and architectures, and exhibits phenomena similar to the "rich" and "lazy" regime. While many network behaviors depend quantitatively on architecture, our findings point to certain behaviors that are widely conserved once models are sufficiently flexible.
title When Representations Align: Universality in Representation Learning Dynamics
topic Machine Learning
Neurons and Cognition
url https://arxiv.org/abs/2402.09142