Why Shallow Networks Struggle to Approximate and Learn High Frequencies

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Shijun, Zhao, Hongkai, Zhong, Yimin, Zhou, Haomin
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910980890427392
author Zhang, Shijun
Zhao, Hongkai
Zhong, Yimin
Zhou, Haomin
author_facet Zhang, Shijun
Zhao, Hongkai
Zhong, Yimin
Zhou, Haomin
contents In this work, we present a comprehensive study combining mathematical and computational analysis to explain why a two-layer neural network struggles to handle high frequencies in both approximation and learning, especially when machine precision, numerical noise, and computational cost are significant factors in practice. Specifically, we investigate the following fundamental computational issues: (1) the minimal numerical error achievable under finite precision, (2) the computational cost required to attain a given accuracy, and (3) the stability of the method with respect to perturbations. The core of our analysis lies in the conditioning of the representation and its learning dynamics. Explicit answers to these questions are provided, along with supporting numerical evidence.
format Preprint
id arxiv_https___arxiv_org_abs_2306_17301
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Why Shallow Networks Struggle to Approximate and Learn High Frequencies
Zhang, Shijun
Zhao, Hongkai
Zhong, Yimin
Zhou, Haomin
Machine Learning
Numerical Analysis
In this work, we present a comprehensive study combining mathematical and computational analysis to explain why a two-layer neural network struggles to handle high frequencies in both approximation and learning, especially when machine precision, numerical noise, and computational cost are significant factors in practice. Specifically, we investigate the following fundamental computational issues: (1) the minimal numerical error achievable under finite precision, (2) the computational cost required to attain a given accuracy, and (3) the stability of the method with respect to perturbations. The core of our analysis lies in the conditioning of the representation and its learning dynamics. Explicit answers to these questions are provided, along with supporting numerical evidence.
title Why Shallow Networks Struggle to Approximate and Learn High Frequencies
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2306.17301