Towards Size-Independent Generalization Bounds for Deep Operator Nets

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gopalani, Pulkit, Karmakar, Sayar, Kumar, Dibyakanti, Mukherjee, Anirbit
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910727214727168
author Gopalani, Pulkit
Karmakar, Sayar
Kumar, Dibyakanti
Mukherjee, Anirbit
author_facet Gopalani, Pulkit
Karmakar, Sayar
Kumar, Dibyakanti
Mukherjee, Anirbit
contents In recent times machine learning methods have made significant advances in becoming a useful tool for analyzing physical systems. A particularly active area in this theme has been "physics-informed machine learning" which focuses on using neural nets for numerically solving differential equations. In this work, we aim to advance the theory of measuring out-of-sample error while training DeepONets - which is among the most versatile ways to solve P.D.E systems in one-shot. Firstly, for a class of DeepONets, we prove a bound on their Rademacher complexity which does not explicitly scale with the width of the nets involved. Secondly, we use this to show how the Huber loss can be chosen so that for these DeepONet classes generalization error bounds can be obtained that have no explicit dependence on the size of the nets. The effective capacity measure for DeepONets that we thus derive is also shown to correlate with the behavior of generalization error in experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2205_11359
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Towards Size-Independent Generalization Bounds for Deep Operator Nets
Gopalani, Pulkit
Karmakar, Sayar
Kumar, Dibyakanti
Mukherjee, Anirbit
Machine Learning
Numerical Analysis
In recent times machine learning methods have made significant advances in becoming a useful tool for analyzing physical systems. A particularly active area in this theme has been "physics-informed machine learning" which focuses on using neural nets for numerically solving differential equations. In this work, we aim to advance the theory of measuring out-of-sample error while training DeepONets - which is among the most versatile ways to solve P.D.E systems in one-shot. Firstly, for a class of DeepONets, we prove a bound on their Rademacher complexity which does not explicitly scale with the width of the nets involved. Secondly, we use this to show how the Huber loss can be chosen so that for these DeepONet classes generalization error bounds can be obtained that have no explicit dependence on the size of the nets. The effective capacity measure for DeepONets that we thus derive is also shown to correlate with the behavior of generalization error in experiments.
title Towards Size-Independent Generalization Bounds for Deep Operator Nets
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2205.11359