Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hu, Yutong, Huang, Quzhe, Tao, Mingxu, Zhang, Chen, Feng, Yansong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917662559305728
author Hu, Yutong
Huang, Quzhe
Tao, Mingxu
Zhang, Chen
Feng, Yansong
author_facet Hu, Yutong
Huang, Quzhe
Tao, Mingxu
Zhang, Chen
Feng, Yansong
contents Recent studies have shown that Large Language Models (LLMs) have the potential to process extremely long text. Many works only evaluate LLMs' long-text processing ability on the language modeling task, with perplexity (PPL) as the evaluation metric. However, in our study, we find that there is no correlation between PPL and LLMs' long-text understanding ability. Besides, PPL may only reflect the model's ability to model local information instead of catching long-range dependency. Therefore, only using PPL to prove the model could process long text is inappropriate. The local focus feature of PPL could also explain some existing phenomena, such as the great extrapolation ability of the position method ALiBi. When evaluating a model's ability in long text, we might pay more attention to PPL's limitation and avoid overly relying on it.
format Preprint
id arxiv_https___arxiv_org_abs_2405_06105
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
Hu, Yutong
Huang, Quzhe
Tao, Mingxu
Zhang, Chen
Feng, Yansong
Computation and Language
Recent studies have shown that Large Language Models (LLMs) have the potential to process extremely long text. Many works only evaluate LLMs' long-text processing ability on the language modeling task, with perplexity (PPL) as the evaluation metric. However, in our study, we find that there is no correlation between PPL and LLMs' long-text understanding ability. Besides, PPL may only reflect the model's ability to model local information instead of catching long-range dependency. Therefore, only using PPL to prove the model could process long text is inappropriate. The local focus feature of PPL could also explain some existing phenomena, such as the great extrapolation ability of the position method ALiBi. When evaluating a model's ability in long text, we might pay more attention to PPL's limitation and avoid overly relying on it.
title Can Perplexity Reflect Large Language Model's Ability in Long Text Understanding?
topic Computation and Language
url https://arxiv.org/abs/2405.06105