Is In-Context Universality Enough? MLPs are Also Universal In-Context
Fuente:
arXiv
Salvato in:
| Autori principali: | , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866929699114975232 |
|---|---|
| author | Kratsios, Anastasis Furuya, Takashi |
| author_facet | Kratsios, Anastasis Furuya, Takashi |
| contents | The success of transformers is often linked to their ability to perform in-context learning. Recent work shows that transformers are universal in context, capable of approximating any real-valued continuous function of a context (a probability measure over $\mathcal{X}\subseteq \mathbb{R}^d$) and a query $x\in \mathcal{X}$. This raises the question: Does in-context universality explain their advantage over classical models? We answer this in the negative by proving that MLPs with trainable activation functions are also universal in-context. This suggests the transformer's success is likely due to other factors like inductive bias or training stability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2502_03327 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Is In-Context Universality Enough? MLPs are Also Universal In-Context Kratsios, Anastasis Furuya, Takashi Machine Learning Numerical Analysis Neural and Evolutionary Computing Probability The success of transformers is often linked to their ability to perform in-context learning. Recent work shows that transformers are universal in context, capable of approximating any real-valued continuous function of a context (a probability measure over $\mathcal{X}\subseteq \mathbb{R}^d$) and a query $x\in \mathcal{X}$. This raises the question: Does in-context universality explain their advantage over classical models? We answer this in the negative by proving that MLPs with trainable activation functions are also universal in-context. This suggests the transformer's success is likely due to other factors like inductive bias or training stability. |
| title | Is In-Context Universality Enough? MLPs are Also Universal In-Context |
| topic | Machine Learning Numerical Analysis Neural and Evolutionary Computing Probability |
| url | https://arxiv.org/abs/2502.03327 |