Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912767386058752 |
|---|---|
| author | Bansal, Rachit Zhang, Aston Tiwari, Rishabh Madaan, Lovish Duvvuri, Sai Surya Khatri, Devvrit Brandfonbrener, David Alvarez-Melis, David Bhargava, Prajjwal Kale, Mihir Sanjay Jelassi, Samy |
| author_facet | Bansal, Rachit Zhang, Aston Tiwari, Rishabh Madaan, Lovish Duvvuri, Sai Surya Khatri, Devvrit Brandfonbrener, David Alvarez-Melis, David Bhargava, Prajjwal Kale, Mihir Sanjay Jelassi, Samy |
| contents | Progress on training and architecture strategies has enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other hand, it has been shown that inference-time compute can be used to scale performance of LLMs, often by generating thinking tokens, on challenging tasks involving multi-step reasoning. Through controlled experiments on sandbox long-context tasks, we find that such inference-time strategies show rapidly diminishing returns and fail at long context. We attribute these failures to score dilution, a phenomenon inherent to static self-attention. Further, we show that current inference-time strategies cannot retrieve relevant long-context signals under certain conditions. We propose a simple method that, through targeted gradient updates on the given context, provably overcomes limitations of static self-attention. We find that this shift in how inference-time compute is spent leads to consistently large performance improvements across models and long-context benchmarks. Our method leads to large 12.6 and 14.1 percentage point improvements for Qwen3-4B on average across subsets of LongBench-v2 and ZeroScrolls benchmarks. The takeaway is practical: for long context, a small amount of context-specific training is a better use of inference compute than current inference-time scaling strategies like producing more thinking tokens. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_13898 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs Bansal, Rachit Zhang, Aston Tiwari, Rishabh Madaan, Lovish Duvvuri, Sai Surya Khatri, Devvrit Brandfonbrener, David Alvarez-Melis, David Bhargava, Prajjwal Kale, Mihir Sanjay Jelassi, Samy Machine Learning Computation and Language Progress on training and architecture strategies has enabled LLMs with millions of tokens in context length. However, empirical evidence suggests that such long-context LLMs can consume far more text than they can reliably use. On the other hand, it has been shown that inference-time compute can be used to scale performance of LLMs, often by generating thinking tokens, on challenging tasks involving multi-step reasoning. Through controlled experiments on sandbox long-context tasks, we find that such inference-time strategies show rapidly diminishing returns and fail at long context. We attribute these failures to score dilution, a phenomenon inherent to static self-attention. Further, we show that current inference-time strategies cannot retrieve relevant long-context signals under certain conditions. We propose a simple method that, through targeted gradient updates on the given context, provably overcomes limitations of static self-attention. We find that this shift in how inference-time compute is spent leads to consistently large performance improvements across models and long-context benchmarks. Our method leads to large 12.6 and 14.1 percentage point improvements for Qwen3-4B on average across subsets of LongBench-v2 and ZeroScrolls benchmarks. The takeaway is practical: for long context, a small amount of context-specific training is a better use of inference compute than current inference-time scaling strategies like producing more thinking tokens. |
| title | Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs |
| topic | Machine Learning Computation and Language |
| url | https://arxiv.org/abs/2512.13898 |