NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deng, Hexuan, Jiao, Wenxiang, Liu, Xuebo, Zhang, Min, Tu, Zhaopeng
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912089606455296
author Deng, Hexuan
Jiao, Wenxiang
Liu, Xuebo
Zhang, Min
Tu, Zhaopeng
author_facet Deng, Hexuan
Jiao, Wenxiang
Liu, Xuebo
Zhang, Min
Tu, Zhaopeng
contents Despite their remarkable abilities in various tasks, large language models (LLMs) still struggle with real-time information (e.g., new facts and terms) due to the knowledge cutoff in their development process. However, existing benchmarks focus on outdated content and limited fields, facing difficulties in real-time updating and leaving new terms unexplored. To address this problem, we propose an adaptive benchmark, NewTerm, for real-time evaluation of new terms. We design a highly automated construction method to ensure high-quality benchmark construction with minimal human effort, allowing flexible updates for real-time information. Empirical results on various LLMs demonstrate over 20% performance reduction caused by new terms. Additionally, while updates to the knowledge cutoff of LLMs can cover some of the new terms, they are unable to generalize to more distant new terms. We also analyze which types of terms are more challenging and why LLMs struggle with new terms, paving the way for future research. Finally, we construct NewTerm 2022 and 2023 to evaluate the new terms updated each year and will continue updating annually. The benchmark and codes can be found at https://github.com/hexuandeng/NewTerm.
format Preprint
id arxiv_https___arxiv_org_abs_2410_20814
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
Deng, Hexuan
Jiao, Wenxiang
Liu, Xuebo
Zhang, Min
Tu, Zhaopeng
Computation and Language
Despite their remarkable abilities in various tasks, large language models (LLMs) still struggle with real-time information (e.g., new facts and terms) due to the knowledge cutoff in their development process. However, existing benchmarks focus on outdated content and limited fields, facing difficulties in real-time updating and leaving new terms unexplored. To address this problem, we propose an adaptive benchmark, NewTerm, for real-time evaluation of new terms. We design a highly automated construction method to ensure high-quality benchmark construction with minimal human effort, allowing flexible updates for real-time information. Empirical results on various LLMs demonstrate over 20% performance reduction caused by new terms. Additionally, while updates to the knowledge cutoff of LLMs can cover some of the new terms, they are unable to generalize to more distant new terms. We also analyze which types of terms are more challenging and why LLMs struggle with new terms, paving the way for future research. Finally, we construct NewTerm 2022 and 2023 to evaluate the new terms updated each year and will continue updating annually. The benchmark and codes can be found at https://github.com/hexuandeng/NewTerm.
title NewTerm: Benchmarking Real-Time New Terms for Large Language Models with Annual Updates
topic Computation and Language
url https://arxiv.org/abs/2410.20814