Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912970523541504 |
|---|---|
| author | Upasani, Shubhangi Wu, Chen Rainton, Jay Li, Bo Thakker, Urmish Hu, Changran Zhang, Qizheng |
| author_facet | Upasani, Shubhangi Wu, Chen Rainton, Jay Li, Bo Thakker, Urmish Hu, Changran Zhang, Qizheng |
| contents | Test-time adaptation enables large language models (LLMs) to modify their behavior at inference without updating model parameters. A common approach is many-shot prompting, where large numbers of in-context learning (ICL) examples are injected as an input-space test-time update. Although performance can improve as more demonstrations are added, the reliability and limits of this update mechanism remain poorly understood, particularly for open-source models. We present an empirical study of many-shot prompting across tasks and model backbones, analyzing how performance varies with update magnitude, example ordering, and selection policy. We further study Dynamic and Reinforced ICL as alternative test-time update strategies that control which information is injected and how it constrains model behavior. We find that many-shot prompting is effective for structured tasks where demonstrations provide high information gain, but is highly sensitive to selection strategy and often shows limited benefits for open-ended generation tasks. Overall, we characterize the practical limits of prompt-based test-time adaptation and outline when input-space updates are beneficial versus harmful. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_05829 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls Upasani, Shubhangi Wu, Chen Rainton, Jay Li, Bo Thakker, Urmish Hu, Changran Zhang, Qizheng Machine Learning Computation and Language Test-time adaptation enables large language models (LLMs) to modify their behavior at inference without updating model parameters. A common approach is many-shot prompting, where large numbers of in-context learning (ICL) examples are injected as an input-space test-time update. Although performance can improve as more demonstrations are added, the reliability and limits of this update mechanism remain poorly understood, particularly for open-source models. We present an empirical study of many-shot prompting across tasks and model backbones, analyzing how performance varies with update magnitude, example ordering, and selection policy. We further study Dynamic and Reinforced ICL as alternative test-time update strategies that control which information is injected and how it constrains model behavior. We find that many-shot prompting is effective for structured tasks where demonstrations provide high information gain, but is highly sensitive to selection strategy and often shows limited benefits for open-ended generation tasks. Overall, we characterize the practical limits of prompt-based test-time adaptation and outline when input-space updates are beneficial versus harmful. |
| title | Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2603.05829 |