The Pontryagin Maximum Principle for Training Convolutional Neural Networks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hofmann, Sebastian, Borzì, Alfio
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916915584172032
author Hofmann, Sebastian
Borzì, Alfio
author_facet Hofmann, Sebastian
Borzì, Alfio
contents A novel batch sequential quadratic Hamiltonian (bSQH) algorithm for training convolutional neural networks (CNNs) with $L^0$-based regularization is presented. This methodology is based on a discrete-time Pontryagin maximum principle (PMP). It uses forward and backward sweeps together with the layerwise approximate maximization of an augmented Hamiltonian function, where the augmentation parameter is chosen adaptively. A technique for determining this augmentation parameter is proposed, and the loss-reduction and convergence properties of the bSQH algorithm are analysed theoretically and validated numerically. Results of numerical experiments in the context of image classification with a sparsity enforcing $L^0$-based regularizer demonstrate the effectiveness of the proposed method in full-batch and mini-batch modes.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11647
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Pontryagin Maximum Principle for Training Convolutional Neural Networks
Hofmann, Sebastian
Borzì, Alfio
Optimization and Control
68T07, 49M05, 65K10
A novel batch sequential quadratic Hamiltonian (bSQH) algorithm for training convolutional neural networks (CNNs) with $L^0$-based regularization is presented. This methodology is based on a discrete-time Pontryagin maximum principle (PMP). It uses forward and backward sweeps together with the layerwise approximate maximization of an augmented Hamiltonian function, where the augmentation parameter is chosen adaptively. A technique for determining this augmentation parameter is proposed, and the loss-reduction and convergence properties of the bSQH algorithm are analysed theoretically and validated numerically. Results of numerical experiments in the context of image classification with a sparsity enforcing $L^0$-based regularizer demonstrate the effectiveness of the proposed method in full-batch and mini-batch modes.
title The Pontryagin Maximum Principle for Training Convolutional Neural Networks
topic Optimization and Control
68T07, 49M05, 65K10
url https://arxiv.org/abs/2504.11647