A PAC-Bayes oracle inequality for sparse neural networks
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866917188581982208 |
|---|---|
| author | Steffen, Maximilian F. Trabs, Mathias |
| author_facet | Steffen, Maximilian F. Trabs, Mathias |
| contents | We study the Gibbs posterior distribution for sparse deep neural nets in a nonparametric regression setting. The posterior can be accessed via Metropolis-adjusted Langevin algorithms. Using a mixture over uniform priors on sparse sets of network weights, we prove an oracle inequality which shows that the method adapts to the unknown regularity and hierarchical structure of the regression function. The estimator achieves the minimax-optimal rate of convergence (up to a logarithmic factor). |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2204_12392 |
| institution | arXiv |
| publishDate | 2022 |
| record_format | arxiv |
| spellingShingle | A PAC-Bayes oracle inequality for sparse neural networks Steffen, Maximilian F. Trabs, Mathias Statistics Theory Machine Learning 62G08, 62F15, 68T05 We study the Gibbs posterior distribution for sparse deep neural nets in a nonparametric regression setting. The posterior can be accessed via Metropolis-adjusted Langevin algorithms. Using a mixture over uniform priors on sparse sets of network weights, we prove an oracle inequality which shows that the method adapts to the unknown regularity and hierarchical structure of the regression function. The estimator achieves the minimax-optimal rate of convergence (up to a logarithmic factor). |
| title | A PAC-Bayes oracle inequality for sparse neural networks |
| topic | Statistics Theory Machine Learning 62G08, 62F15, 68T05 |
| url | https://arxiv.org/abs/2204.12392 |