Understanding In-Context Learning from Repetitions
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909114151469056 |
|---|---|
| author | Yan, Jianhao Xu, Jin Song, Chiyu Wu, Chenming Li, Yafu Zhang, Yue |
| author_facet | Yan, Jianhao Xu, Jin Song, Chiyu Wu, Chenming Li, Yafu Zhang, Yue |
| contents | This paper explores the elusive mechanism underpinning in-context learning in Large Language Models (LLMs). Our work provides a novel perspective by examining in-context learning via the lens of surface repetitions. We quantitatively investigate the role of surface features in text generation, and empirically establish the existence of \emph{token co-occurrence reinforcement}, a principle that strengthens the relationship between two tokens based on their contextual co-occurrences. By investigating the dual impacts of these features, our research illuminates the internal workings of in-context learning and expounds on the reasons for its failures. This paper provides an essential contribution to the understanding of in-context learning and its potential limitations, providing a fresh perspective on this exciting capability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_00297 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Understanding In-Context Learning from Repetitions Yan, Jianhao Xu, Jin Song, Chiyu Wu, Chenming Li, Yafu Zhang, Yue Computation and Language This paper explores the elusive mechanism underpinning in-context learning in Large Language Models (LLMs). Our work provides a novel perspective by examining in-context learning via the lens of surface repetitions. We quantitatively investigate the role of surface features in text generation, and empirically establish the existence of \emph{token co-occurrence reinforcement}, a principle that strengthens the relationship between two tokens based on their contextual co-occurrences. By investigating the dual impacts of these features, our research illuminates the internal workings of in-context learning and expounds on the reasons for its failures. This paper provides an essential contribution to the understanding of in-context learning and its potential limitations, providing a fresh perspective on this exciting capability. |
| title | Understanding In-Context Learning from Repetitions |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2310.00297 |