Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866914004652261376 |
|---|---|
| author | Yu, Yijiong Huang, Yongfeng Qi, Zhixiao Wang, Wei Liu, Weifeng Chen, Ran Pei, Ji |
| author_facet | Yu, Yijiong Huang, Yongfeng Qi, Zhixiao Wang, Wei Liu, Weifeng Chen, Ran Pei, Ji |
| contents | Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_04422 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps Yu, Yijiong Huang, Yongfeng Qi, Zhixiao Wang, Wei Liu, Weifeng Chen, Ran Pei, Ji Computation and Language Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks. |
| title | Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2410.04422 |