Only Strict Saddles in the Energy Landscape of Predictive Coding Networks?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Innocenti, Francesco, Achour, El Mehdi, Singh, Ryan, Buckley, Christopher L.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912112766353408
author Innocenti, Francesco
Achour, El Mehdi
Singh, Ryan
Buckley, Christopher L.
author_facet Innocenti, Francesco
Achour, El Mehdi
Singh, Ryan
Buckley, Christopher L.
contents Predictive coding (PC) is an energy-based learning algorithm that performs iterative inference over network activities before updating weights. Recent work suggests that PC can converge in fewer learning steps than backpropagation thanks to its inference procedure. However, these advantages are not always observed, and the impact of PC inference on learning is not theoretically well understood. Here, we study the geometry of the PC energy landscape at the inference equilibrium of the network activities. For deep linear networks, we first show that the equilibrated energy is simply a rescaled mean squared error loss with a weight-dependent rescaling. We then prove that many highly degenerate (non-strict) saddles of the loss including the origin become much easier to escape (strict) in the equilibrated energy. Our theory is validated by experiments on both linear and non-linear networks. Based on these and other results, we conjecture that all the saddles of the equilibrated energy are strict. Overall, this work suggests that PC inference makes the loss landscape more benign and robust to vanishing gradients, while also highlighting the fundamental challenge of scaling PC to deeper models.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11979
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Only Strict Saddles in the Energy Landscape of Predictive Coding Networks?
Innocenti, Francesco
Achour, El Mehdi
Singh, Ryan
Buckley, Christopher L.
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
I.2.6
Predictive coding (PC) is an energy-based learning algorithm that performs iterative inference over network activities before updating weights. Recent work suggests that PC can converge in fewer learning steps than backpropagation thanks to its inference procedure. However, these advantages are not always observed, and the impact of PC inference on learning is not theoretically well understood. Here, we study the geometry of the PC energy landscape at the inference equilibrium of the network activities. For deep linear networks, we first show that the equilibrated energy is simply a rescaled mean squared error loss with a weight-dependent rescaling. We then prove that many highly degenerate (non-strict) saddles of the loss including the origin become much easier to escape (strict) in the equilibrated energy. Our theory is validated by experiments on both linear and non-linear networks. Based on these and other results, we conjecture that all the saddles of the equilibrated energy are strict. Overall, this work suggests that PC inference makes the loss landscape more benign and robust to vanishing gradients, while also highlighting the fundamental challenge of scaling PC to deeper models.
title Only Strict Saddles in the Energy Landscape of Predictive Coding Networks?
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
I.2.6
url https://arxiv.org/abs/2408.11979