Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishikawa, Satoki, Yokota, Rio, Karakida, Ryo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918027366236160
author Ishikawa, Satoki
Yokota, Rio
Karakida, Ryo
author_facet Ishikawa, Satoki
Yokota, Rio
Karakida, Ryo
contents Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters because of the locality, making it challenging to identify desirable settings in which the algorithm progresses in a stable manner. To provide theoretical and quantitative insights, we introduce the maximal update parameterization ($μ$P) in the infinite-width limit for two representative designs of local targets: predictive coding (PC) and target propagation (TP). We verified that $μ$P enables hyperparameter transfer across models of different widths. Furthermore, our analysis revealed unique and intriguing properties of $μ$P that are not present in conventional BP. By analyzing deep linear networks, we found that PC's gradients interpolate between first-order and Gauss-Newton-like gradients, depending on the parameterization. We demonstrate that, in specific standard settings, PC in the infinite-width limit behaves more similarly to the first-order gradient. For TP, even with the standard scaling of the last layer, which differs from classical $μ$P, its local loss optimization favors the feature learning regime over the kernel regime.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02001
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
Ishikawa, Satoki
Yokota, Rio
Karakida, Ryo
Machine Learning
Neural and Evolutionary Computing
Local learning, which trains a network through layer-wise local targets and losses, has been studied as an alternative to backpropagation (BP) in neural computation. However, its algorithms often become more complex or require additional hyperparameters because of the locality, making it challenging to identify desirable settings in which the algorithm progresses in a stable manner. To provide theoretical and quantitative insights, we introduce the maximal update parameterization ($μ$P) in the infinite-width limit for two representative designs of local targets: predictive coding (PC) and target propagation (TP). We verified that $μ$P enables hyperparameter transfer across models of different widths. Furthermore, our analysis revealed unique and intriguing properties of $μ$P that are not present in conventional BP. By analyzing deep linear networks, we found that PC's gradients interpolate between first-order and Gauss-Newton-like gradients, depending on the parameterization. We demonstrate that, in specific standard settings, PC in the infinite-width limit behaves more similarly to the first-order gradient. For TP, even with the standard scaling of the last layer, which differs from classical $μ$P, its local loss optimization favors the feature learning regime over the kernel regime.
title Local Loss Optimization in the Infinite Width: Stable Parameterization of Predictive Coding Networks and Target Propagation
topic Machine Learning
Neural and Evolutionary Computing
url https://arxiv.org/abs/2411.02001