Marathon: A Race Through the Realm of Long Context with Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhang, Lei, Li, Yunshui, Liu, Ziqiang, yang, Jiaxi, Liu, Junhao, Chen, Longze, Luo, Run, Yang, Min
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916300808257536
author Zhang, Lei
Li, Yunshui
Liu, Ziqiang
yang, Jiaxi
Liu, Junhao
Chen, Longze
Luo, Run
Yang, Min
author_facet Zhang, Lei
Li, Yunshui
Liu, Ziqiang
yang, Jiaxi
Liu, Junhao
Chen, Longze
Luo, Run
Yang, Min
contents With the advancement of large language models (LLMs) and the expansion of their context windows, existing long-context benchmarks fall short in effectively evaluating the models' comprehension and reasoning abilities in extended texts. Moreover, conventional benchmarks relying on F1 metrics often inaccurately score responses: they may undervalue correct answers that differ from the reference responses and overvalue incorrect ones that resemble the reference texts. In response to these limitations, we introduce Marathon, a novel evaluation benchmark that adopts a multiple-choice question format. It is specifically designed to overcome the constraints of previous benchmarks and provide a rapid, precise, and unbiased appraisal of the long-context comprehension skills of large language models. We conducted comprehensive evaluations on the Marathon benchmark with a range of state-of-the-art LLMs and assessed the effectiveness of various optimization strategies tailored for long-context generation. We anticipate that the Marathon benchmark and its associated leaderboard will enable a more precise and equitable evaluation of LLMs' capabilities in understanding and reasoning over extended contexts. Marathon is available at https://github.com/Hambaobao/Marathon.
format Preprint
id arxiv_https___arxiv_org_abs_2312_09542
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Marathon: A Race Through the Realm of Long Context with Large Language Models
Zhang, Lei
Li, Yunshui
Liu, Ziqiang
yang, Jiaxi
Liu, Junhao
Chen, Longze
Luo, Run
Yang, Min
Computation and Language
With the advancement of large language models (LLMs) and the expansion of their context windows, existing long-context benchmarks fall short in effectively evaluating the models' comprehension and reasoning abilities in extended texts. Moreover, conventional benchmarks relying on F1 metrics often inaccurately score responses: they may undervalue correct answers that differ from the reference responses and overvalue incorrect ones that resemble the reference texts. In response to these limitations, we introduce Marathon, a novel evaluation benchmark that adopts a multiple-choice question format. It is specifically designed to overcome the constraints of previous benchmarks and provide a rapid, precise, and unbiased appraisal of the long-context comprehension skills of large language models. We conducted comprehensive evaluations on the Marathon benchmark with a range of state-of-the-art LLMs and assessed the effectiveness of various optimization strategies tailored for long-context generation. We anticipate that the Marathon benchmark and its associated leaderboard will enable a more precise and equitable evaluation of LLMs' capabilities in understanding and reasoning over extended contexts. Marathon is available at https://github.com/Hambaobao/Marathon.
title Marathon: A Race Through the Realm of Long Context with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2312.09542