Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bansal, Rachit, Zhang, Aston, Tiwari, Rishabh, Madaan, Lovish, Duvvuri, Sai Surya, Khatri, Devvrit, Brandfonbrener, David, Alvarez-Melis, David, Bhargava, Prajjwal, Kale, Mihir Sanjay, Jelassi, Samy
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912767386058752
author Bansal, Rachit
Zhang, Aston
Tiwari, Rishabh
Madaan, Lovish
Duvvuri, Sai Surya
Khatri, Devvrit
Brandfonbrener, David
Alvarez-Melis, David
Bhargava, Prajjwal
Kale, Mihir Sanjay
Jelassi, Samy
author_facet Bansal, Rachit
Zhang, Aston
Tiwari, Rishabh
Madaan, Lovish
Duvvuri, Sai Surya
Khatri, Devvrit
Brandfonbrener, David
Alvarez-Melis, David
Bhargava, Prajjwal
Kale, Mihir Sanjay
Jelassi, Samy
contents Progress on training and architecture strategies has enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other hand, it has been shown that inference-time compute can be used to scale performance of LLMs, often by generating thinking tokens, on challenging tasks involving multi-step reasoning. Through controlled experiments on sandbox long-context tasks, we find that such inference-time strategies show rapidly diminishing returns and fail at long context. We attribute these failures to score dilution, a phenomenon inherent to static self-attention. Further, we show that current inference-time strategies cannot retrieve relevant long-context signals under certain conditions. We propose a simple method that, through targeted gradient updates on the given context, provably overcomes limitations of static self-attention. We find that this shift in how inference-time compute is spent leads to consistently large performance improvements across models and long-context benchmarks. Our method leads to large 12.6 and 14.1 percentage point improvements for Qwen3-4B on average across subsets of LongBench-v2 and ZeroScrolls benchmarks. The takeaway is practical: for long context, a small amount of context-specific training is a better use of inference compute than current inference-time scaling strategies like producing more thinking tokens.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13898
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
Bansal, Rachit
Zhang, Aston
Tiwari, Rishabh
Madaan, Lovish
Duvvuri, Sai Surya
Khatri, Devvrit
Brandfonbrener, David
Alvarez-Melis, David
Bhargava, Prajjwal
Kale, Mihir Sanjay
Jelassi, Samy
Machine Learning
Computation and Language
Progress on training and architecture strategies has enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other hand, it has been shown that inference-time compute can be used to scale performance of LLMs, often by generating thinking tokens, on challenging tasks involving multi-step reasoning. Through controlled experiments on sandbox long-context tasks, we find that such inference-time strategies show rapidly diminishing returns and fail at long context. We attribute these failures to score dilution, a phenomenon inherent to static self-attention. Further, we show that current inference-time strategies cannot retrieve relevant long-context signals under certain conditions. We propose a simple method that, through targeted gradient updates on the given context, provably overcomes limitations of static self-attention. We find that this shift in how inference-time compute is spent leads to consistently large performance improvements across models and long-context benchmarks. Our method leads to large 12.6 and 14.1 percentage point improvements for Qwen3-4B on average across subsets of LongBench-v2 and ZeroScrolls benchmarks. The takeaway is practical: for long context, a small amount of context-specific training is a better use of inference compute than current inference-time scaling strategies like producing more thinking tokens.
title Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2512.13898