ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lee, Taewhoo, Yoon, Chanwoong, Jang, Kyochul, Lee, Donghyeon, Song, Minju, Kim, Hyunjae, Kang, Jaewoo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910847031312384
author Lee, Taewhoo
Yoon, Chanwoong
Jang, Kyochul
Lee, Donghyeon
Song, Minju
Kim, Hyunjae
Kang, Jaewoo
author_facet Lee, Taewhoo
Yoon, Chanwoong
Jang, Kyochul
Lee, Donghyeon
Song, Minju
Kim, Hyunjae
Kang, Jaewoo
contents Recent advancements in large language models (LLM) capable of processing extremely long texts highlight the need for a dedicated evaluation benchmark to assess their long-context capabilities. However, existing methods, like the needle-in-a-haystack test, do not effectively assess whether these models fully utilize contextual information, raising concerns about the reliability of current evaluation techniques. To thoroughly examine the effectiveness of existing benchmarks, we introduce a new metric called information coverage (IC), which quantifies the proportion of the input context necessary for answering queries. Our findings indicate that current benchmarks exhibit low IC; although the input context may be extensive, the actual usable context is often limited. To address this, we present ETHIC, a novel benchmark designed to assess LLMs' ability to leverage the entire context. Our benchmark comprises 1,986 test instances spanning four long-context tasks with high IC scores in the domains of books, debates, medicine, and law. Our evaluations reveal significant performance drops in contemporary LLMs, highlighting a critical challenge in managing long contexts. Our benchmark is available at https://github.com/dmis-lab/ETHIC.
format Preprint
id arxiv_https___arxiv_org_abs_2410_16848
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
Lee, Taewhoo
Yoon, Chanwoong
Jang, Kyochul
Lee, Donghyeon
Song, Minju
Kim, Hyunjae
Kang, Jaewoo
Computation and Language
Recent advancements in large language models (LLM) capable of processing extremely long texts highlight the need for a dedicated evaluation benchmark to assess their long-context capabilities. However, existing methods, like the needle-in-a-haystack test, do not effectively assess whether these models fully utilize contextual information, raising concerns about the reliability of current evaluation techniques. To thoroughly examine the effectiveness of existing benchmarks, we introduce a new metric called information coverage (IC), which quantifies the proportion of the input context necessary for answering queries. Our findings indicate that current benchmarks exhibit low IC; although the input context may be extensive, the actual usable context is often limited. To address this, we present ETHIC, a novel benchmark designed to assess LLMs' ability to leverage the entire context. Our benchmark comprises 1,986 test instances spanning four long-context tasks with high IC scores in the domains of books, debates, medicine, and law. Our evaluations reveal significant performance drops in contemporary LLMs, highlighting a critical challenge in managing long contexts. Our benchmark is available at https://github.com/dmis-lab/ETHIC.
title ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
topic Computation and Language
url https://arxiv.org/abs/2410.16848