Saved in:
Bibliographic Details
Main Authors: Liu, Xinyu, Zhao, Runsong, Huang, Pengcheng, Xiao, Chunyang, Li, Bei, Wang, Jingang, Xiao, Tong, Zhu, Jingbo
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2410.04727
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912061975429120
author Liu, Xinyu
Zhao, Runsong
Huang, Pengcheng
Xiao, Chunyang
Li, Bei
Wang, Jingang
Xiao, Tong
Zhu, Jingbo
author_facet Liu, Xinyu
Zhao, Runsong
Huang, Pengcheng
Xiao, Chunyang
Li, Bei
Wang, Jingang
Xiao, Tong
Zhu, Jingbo
contents Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model's effective memorization length. However, through thorough investigations, we find limitations for currently existing evaluations on model's memorization capability. We provide an extensive survey for limitations in this work and propose a new method called forgetting curve to measure the memorization capability of long-context models. We show that forgetting curve has the advantage of being robust to the tested corpus and the experimental settings, of not relying on prompts and can be applied to any model size. We apply our forgetting curve to a large variety of models involving both transformer and RNN/SSM based architectures. Our measurement provides empirical evidence for the effectiveness of transformer extension techniques while raises questions for the effective length of RNN/SSM based models. We also examine the difference between our measurement and existing benchmarks as well as popular metrics for various models. Our code and results can be found at https://github.com/1azybug/ForgettingCurve.
format Preprint
id arxiv_https___arxiv_org_abs_2410_04727
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
Liu, Xinyu
Zhao, Runsong
Huang, Pengcheng
Xiao, Chunyang
Li, Bei
Wang, Jingang
Xiao, Tong
Zhu, Jingbo
Computation and Language
Numerous recent works target to extend effective context length for language models and various methods, tasks and benchmarks exist to measure model's effective memorization length. However, through thorough investigations, we find limitations for currently existing evaluations on model's memorization capability. We provide an extensive survey for limitations in this work and propose a new method called forgetting curve to measure the memorization capability of long-context models. We show that forgetting curve has the advantage of being robust to the tested corpus and the experimental settings, of not relying on prompts and can be applied to any model size. We apply our forgetting curve to a large variety of models involving both transformer and RNN/SSM based architectures. Our measurement provides empirical evidence for the effectiveness of transformer extension techniques while raises questions for the effective length of RNN/SSM based models. We also examine the difference between our measurement and existing benchmarks as well as popular metrics for various models. Our code and results can be found at https://github.com/1azybug/ForgettingCurve.
title Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-context Models
topic Computation and Language
url https://arxiv.org/abs/2410.04727