Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yu, Yijiong, Huang, Yongfeng, Qi, Zhixiao, Wang, Wei, Liu, Weifeng, Chen, Ran, Pei, Ji
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914004652261376
author Yu, Yijiong
Huang, Yongfeng
Qi, Zhixiao
Wang, Wei
Liu, Weifeng
Chen, Ran
Pei, Ji
author_facet Yu, Yijiong
Huang, Yongfeng
Qi, Zhixiao
Wang, Wei
Liu, Weifeng
Chen, Ran
Pei, Ji
contents Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04422
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
Yu, Yijiong
Huang, Yongfeng
Qi, Zhixiao
Wang, Wei
Liu, Weifeng
Chen, Ran
Pei, Ji
Computation and Language
Long-context language models (LCLMs), characterized by their extensive context window, are becoming popular. However, despite the fact that they are nearly perfect at standard long-context retrieval tasks, our evaluations demonstrate they fail in some basic cases. Later, we find they can be well addressed with a sufficient number of reasoning steps, guided by specific CoT prompts. This result emphasizes the potential necessity of solving specific long-context tasks using long-CoT methods, while previous long-context benchmarks always ignore the necessity of long reasoning for long-context tasks and treat them as direct QA tasks.
title Long-context Language Models Fail in Basic Retrieval Tasks Without Sufficient Reasoning Steps
topic Computation and Language
url https://arxiv.org/abs/2410.04422