Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hamilton, Sil, Hicke, Rebecca M. M., Wilkens, Matthew, Mimno, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908372907851776
author Hamilton, Sil
Hicke, Rebecca M. M.
Wilkens, Matthew
Mimno, David
author_facet Hamilton, Sil
Hicke, Rebecca M. M.
Wilkens, Matthew
Mimno, David
contents Although the context length of large language models (LLMs) has increased to millions of tokens, evaluating their effectiveness beyond needle-in-a-haystack approaches has proven difficult. We argue that novels provide a case study of subtle, complicated structure and long-range semantic dependencies often over 128k tokens in length. Inspired by work on computational novel analysis, we release the Too Long, Didn't Model (TLDM) benchmark, which tests a model's ability to report plot summary, storyworld configuration, and elapsed narrative time. We find that none of seven tested frontier LLMs retain stable understanding beyond 64k tokens. Our results suggest language model developers must look beyond "lost in the middle" benchmarks when evaluating model performance in complex long-context scenarios. To aid in further development we release the TLDM benchmark together with reference code and data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_14925
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
Hamilton, Sil
Hicke, Rebecca M. M.
Wilkens, Matthew
Mimno, David
Computation and Language
Artificial Intelligence
Machine Learning
Although the context length of large language models (LLMs) has increased to millions of tokens, evaluating their effectiveness beyond needle-in-a-haystack approaches has proven difficult. We argue that novels provide a case study of subtle, complicated structure and long-range semantic dependencies often over 128k tokens in length. Inspired by work on computational novel analysis, we release the Too Long, Didn't Model (TLDM) benchmark, which tests a model's ability to report plot summary, storyworld configuration, and elapsed narrative time. We find that none of seven tested frontier LLMs retain stable understanding beyond 64k tokens. Our results suggest language model developers must look beyond "lost in the middle" benchmarks when evaluating model performance in complex long-context scenarios. To aid in further development we release the TLDM benchmark together with reference code and data.
title Too Long, Didn't Model: Decomposing LLM Long-Context Understanding With Novels
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.14925