Training Neural Networks with Optimal Double-Bayesian Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Bui, Vy, Yu, Hang, Kantipudi, Karthik, Yaniv, Ziv, Jaeger, Stefan
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917511363035136
author Bui, Vy
Yu, Hang
Kantipudi, Karthik
Yaniv, Ziv
Jaeger, Stefan
author_facet Bui, Vy
Yu, Hang
Kantipudi, Karthik
Yaniv, Ziv
Jaeger, Stefan
contents Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, finding optimal hyperparameters to guide training has proven challenging. While it is widely acknowledged that selecting appropriate parameters is crucial for avoiding overfitting and achieving unbiased outcomes, this choice remains largely based on empirical experiments and experience. This paper presents a new probabilistic framework for the learning rate, a key parameter in stochastic gradient descent. The framework develops classic Bayesian statistics into a double-Bayesian decision mechanism involving two antagonistic Bayesian processes. A theoretically optimal learning rate can be derived from these two processes and used for stochastic gradient descent. Experiments across various classification, segmentation, and detection tasks corroborate the practical significance of the theoretically derived learning rate. The paper also discusses the ramifications of the proposed double-Bayesian framework for network training and model performance.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20009
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Training Neural Networks with Optimal Double-Bayesian Learning
Bui, Vy
Yu, Hang
Kantipudi, Karthik
Yaniv, Ziv
Jaeger, Stefan
Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, finding optimal hyperparameters to guide training has proven challenging. While it is widely acknowledged that selecting appropriate parameters is crucial for avoiding overfitting and achieving unbiased outcomes, this choice remains largely based on empirical experiments and experience. This paper presents a new probabilistic framework for the learning rate, a key parameter in stochastic gradient descent. The framework develops classic Bayesian statistics into a double-Bayesian decision mechanism involving two antagonistic Bayesian processes. A theoretically optimal learning rate can be derived from these two processes and used for stochastic gradient descent. Experiments across various classification, segmentation, and detection tasks corroborate the practical significance of the theoretically derived learning rate. The paper also discusses the ramifications of the proposed double-Bayesian framework for network training and model performance.
title Training Neural Networks with Optimal Double-Bayesian Learning
topic Machine Learning
Artificial Intelligence
Neural and Evolutionary Computing
url https://arxiv.org/abs/2605.20009