Saved in:
Bibliographic Details
Main Authors: Lu, Taiming, Gao, Muhan, Yu, Kuai, Byerly, Adam, Khashabi, Daniel
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2406.14673
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910633981640704
author Lu, Taiming
Gao, Muhan
Yu, Kuai
Byerly, Adam
Khashabi, Daniel
author_facet Lu, Taiming
Gao, Muhan
Yu, Kuai
Byerly, Adam
Khashabi, Daniel
contents Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts. Our study explores LLMs' long-context reasoning by probing their hidden representations. We find that while LLMs encode the position of target information, they often fail to leverage this in generating accurate responses. This reveals a disconnect between information retrieval and utilization, a "know but don't tell" phenomenon. We further analyze the relationship between extraction time and final accuracy, offering insights into the underlying mechanics of transformer models.
format Preprint
id arxiv_https___arxiv_org_abs_2406_14673
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
Lu, Taiming
Gao, Muhan
Yu, Kuai
Byerly, Adam
Khashabi, Daniel
Computation and Language
Large Language Models (LLMs) exhibit positional bias, struggling to utilize information from the middle or end of long contexts. Our study explores LLMs' long-context reasoning by probing their hidden representations. We find that while LLMs encode the position of target information, they often fail to leverage this in generating accurate responses. This reveals a disconnect between information retrieval and utilization, a "know but don't tell" phenomenon. We further analyze the relationship between extraction time and final accuracy, offering insights into the underlying mechanics of transformer models.
title Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
topic Computation and Language
url https://arxiv.org/abs/2406.14673