Why Shallow Networks Struggle to Approximate and Learn High Frequencies
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910980890427392 |
|---|---|
| author | Zhang, Shijun Zhao, Hongkai Zhong, Yimin Zhou, Haomin |
| author_facet | Zhang, Shijun Zhao, Hongkai Zhong, Yimin Zhou, Haomin |
| contents | In this work, we present a comprehensive study combining mathematical and computational analysis to explain why a two-layer neural network struggles to handle high frequencies in both approximation and learning, especially when machine precision, numerical noise, and computational cost are significant factors in practice. Specifically, we investigate the following fundamental computational issues: (1) the minimal numerical error achievable under finite precision, (2) the computational cost required to attain a given accuracy, and (3) the stability of the method with respect to perturbations. The core of our analysis lies in the conditioning of the representation and its learning dynamics. Explicit answers to these questions are provided, along with supporting numerical evidence. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2306_17301 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Why Shallow Networks Struggle to Approximate and Learn High Frequencies Zhang, Shijun Zhao, Hongkai Zhong, Yimin Zhou, Haomin Machine Learning Numerical Analysis In this work, we present a comprehensive study combining mathematical and computational analysis to explain why a two-layer neural network struggles to handle high frequencies in both approximation and learning, especially when machine precision, numerical noise, and computational cost are significant factors in practice. Specifically, we investigate the following fundamental computational issues: (1) the minimal numerical error achievable under finite precision, (2) the computational cost required to attain a given accuracy, and (3) the stability of the method with respect to perturbations. The core of our analysis lies in the conditioning of the representation and its learning dynamics. Explicit answers to these questions are provided, along with supporting numerical evidence. |
| title | Why Shallow Networks Struggle to Approximate and Learn High Frequencies |
| topic | Machine Learning Numerical Analysis |
| url | https://arxiv.org/abs/2306.17301 |