Mixture of Experts Softens the Curse of Dimensionality in Operator Learning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Kratsios, Anastasis, Furuya, Takashi, Benitez, Jose Antonio Lara, Lassas, Matti, de Hoop, Maarten
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909938974982144
author Kratsios, Anastasis
Furuya, Takashi
Benitez, Jose Antonio Lara
Lassas, Matti
de Hoop, Maarten
author_facet Kratsios, Anastasis
Furuya, Takashi
Benitez, Jose Antonio Lara
Lassas, Matti
de Hoop, Maarten
contents We study the approximation-theoretic implications of mixture-of-experts architectures for operator learning, where the complexity of a single large neural operator is distributed across many small neural operators (NOs), and each input is routed to exactly one NO via a decision tree. We analyze how this tree-based routing and expert decomposition affect approximation power, sample complexity, and stability. Our main result is a distributed universal approximation theorem for mixture of neural operators (MoNOs): any Lipschitz nonlinear operator between $L^2([0,1]^d)$ spaces can be uniformly approximated over the Sobolev unit ball to arbitrary accuracy $\varepsilon>0$ by an MoNO, where each expert NO has a depth, width, and rank scaling as $\mathcal{O}(\varepsilon^{-1})$. Although the number of experts may grow with accuracy, each NO remains small, enough to fit within active memory of standard hardware for reasonable accuracy levels. Our analysis also yields new quantitative approximation rates for classical NOs approximating uniformly continuous nonlinear operators uniformly on compact subsets of $L^2([0,1]^d)$.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09101
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mixture of Experts Softens the Curse of Dimensionality in Operator Learning
Kratsios, Anastasis
Furuya, Takashi
Benitez, Jose Antonio Lara
Lassas, Matti
de Hoop, Maarten
Machine Learning
Artificial Intelligence
Numerical Analysis
We study the approximation-theoretic implications of mixture-of-experts architectures for operator learning, where the complexity of a single large neural operator is distributed across many small neural operators (NOs), and each input is routed to exactly one NO via a decision tree. We analyze how this tree-based routing and expert decomposition affect approximation power, sample complexity, and stability. Our main result is a distributed universal approximation theorem for mixture of neural operators (MoNOs): any Lipschitz nonlinear operator between $L^2([0,1]^d)$ spaces can be uniformly approximated over the Sobolev unit ball to arbitrary accuracy $\varepsilon>0$ by an MoNO, where each expert NO has a depth, width, and rank scaling as $\mathcal{O}(\varepsilon^{-1})$. Although the number of experts may grow with accuracy, each NO remains small, enough to fit within active memory of standard hardware for reasonable accuracy levels. Our analysis also yields new quantitative approximation rates for classical NOs approximating uniformly continuous nonlinear operators uniformly on compact subsets of $L^2([0,1]^d)$.
title Mixture of Experts Softens the Curse of Dimensionality in Operator Learning
topic Machine Learning
Artificial Intelligence
Numerical Analysis
url https://arxiv.org/abs/2404.09101