Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Bansal, Rachit, Zhang, Aston, Tiwari, Rishabh, Madaan, Lovish, Duvvuri, Sai Surya, Khatri, Devvrit, Brandfonbrener, David, Alvarez-Melis, David, Bhargava, Prajjwal, Kale, Mihir Sanjay, Jelassi, Samy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
Interleaved Head Attention
by: Duvvuri, Sai Surya, et al.
Published: (2026)
by: Duvvuri, Sai Surya, et al.
Published: (2026)
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024)
by: Hou, Kaiying, et al.
Published: (2024)
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)
by: Golowich, Noah, et al.
Published: (2025)
Q-Probe: A Lightweight Approach to Reward Maximization for Language Models
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
To Backtrack or Not to Backtrack: When Sequential Search Limits Model Reasoning
by: Qin, Tian, et al.
Published: (2025)
by: Qin, Tian, et al.
Published: (2025)
Mixture of Parrots: Experts improve memorization more than reasoning
by: Jelassi, Samy, et al.
Published: (2024)
by: Jelassi, Samy, et al.
Published: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
by: Mirtaheri, Parsa, et al.
Published: (2025)
by: Mirtaheri, Parsa, et al.
Published: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
Lost in Inference: Rediscovering the Role of Natural Language Inference for Large Language Models
by: Madaan, Lovish, et al.
Published: (2024)
by: Madaan, Lovish, et al.
Published: (2024)
A history of advocacy in breast cancer: lessons for all of us in putting advocacy to work – or ‘how to just get things done!’
by: Christobel M. Saunders, et al.
Published: (2024)
by: Christobel M. Saunders, et al.
Published: (2024)
Collective Model Intelligence Requires Compatible Specialization
by: Pari, Jyothish, et al.
Published: (2024)
by: Pari, Jyothish, et al.
Published: (2024)
Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
by: Tang, Yunhao, et al.
Published: (2025)
by: Tang, Yunhao, et al.
Published: (2025)
Compressing Many-Shots in In-Context Learning
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
LASER: Attention with Exponential Transformation
by: Duvvuri, Sai Surya, et al.
Published: (2024)
by: Duvvuri, Sai Surya, et al.
Published: (2024)
How Does Overparameterization Affect Features?
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
by: Duzgun, Ahmet Cagri, et al.
Published: (2024)
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
by: Liu, Bingbin, et al.
Published: (2025)
by: Liu, Bingbin, et al.
Published: (2025)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
HARP: A challenging human-annotated math reasoning benchmark
by: Yue, Albert S., et al.
Published: (2024)
by: Yue, Albert S., et al.
Published: (2024)
Dual-Encoders for Extreme Multi-Label Classification
by: Gupta, Nilesh, et al.
Published: (2023)
by: Gupta, Nilesh, et al.
Published: (2023)
Making Food Supply Chains Circular in an Emerging Economy: A Non‐Parametric Analysis of Barriers and Strategies in the Indian Context
by: Lovish Raheja, et al.
Published: (2025)
by: Lovish Raheja, et al.
Published: (2025)
In-Context Learning with Iterative Demonstration Selection
by: Qin, Chengwei, et al.
Published: (2023)
by: Qin, Chengwei, et al.
Published: (2023)
When should we prefer Decision Transformers for Offline Reinforcement Learning?
by: Bhargava, Prajjwal, et al.
Published: (2023)
by: Bhargava, Prajjwal, et al.
Published: (2023)
Let's put the clues together and find the aetiology in this patient with dilated right heart!
by: Duygu Inan, et al.
Published: (2024)
by: Duygu Inan, et al.
Published: (2024)
Measures of Information Reflect Memorization Patterns
by: Bansal, Rachit, et al.
Published: (2022)
by: Bansal, Rachit, et al.
Published: (2022)
First record of Celaenorrhinus ratna daphne Evans, 1949 from Himachal Pradesh and its first photographic record from the Western Himalayas (Lepidoptera: Hesperiidae, Pyrginae)
by: Lovish Garlani
Published: (2022)
by: Lovish Garlani
Published: (2022)
Unveiling the Hidden Gem: An Observational Report, Taxonomic Insights and First Photographic Evidence of Pseudochazara baldiva Moore, 1865, from India (Lepidoptera: Nymphalidae)
by: Lovish Garlani
Published: (2024)
by: Lovish Garlani
Published: (2024)
A detailed study of the variations found in the chrysalises of Aglais caschmirensis Kollar, 1844 (Lepidoptera: Papilionoidea, Nymphalidae)
by: Lovish Garlani
Published: (2023)
by: Lovish Garlani
Published: (2023)
Annotated Checklist of Rhopalocera of Himachal Pradesh, India (Insecta: Lepidoptera)
by: Lovish Garlan
Published: (2024)
by: Lovish Garlan
Published: (2024)
Context features influence alcohol reward and motivation
by: Surya Pandey, et al.
Published: (2026)
by: Surya Pandey, et al.
Published: (2026)
LUCID: Attention with Preconditioned Representations
by: Duvvuri, Sai Surya, et al.
Published: (2026)
by: Duvvuri, Sai Surya, et al.
Published: (2026)
Are Emergent Abilities in Large Language Models just In-Context Learning?
by: Lu, Sheng, et al.
Published: (2023)
by: Lu, Sheng, et al.
Published: (2023)
Context-Free Synthetic Data Mitigates Forgetting
by: Bansal, Parikshit, et al.
Published: (2025)
by: Bansal, Parikshit, et al.
Published: (2025)
In-Context Principle Learning from Mistakes
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
by: Zhao, Rosie, et al.
Published: (2025)
by: Zhao, Rosie, et al.
Published: (2025)
You can just review things: A digital ethnography of informal peer review
by: Patel, Jay, et al.
Published: (2026)
by: Patel, Jay, et al.
Published: (2026)
GQ-VAE: A gated quantized VAE for learning variable length tokens
by: Datta, Theo, et al.
Published: (2025)
by: Datta, Theo, et al.
Published: (2025)
“To put an end to this damned thing”: Rebutting denialism strategies performed by people in situation of homelessness during the COVID-19 pandemics (Brazil)
by: Ana Gretel Echazú Böschemeier
Published: (2023)
by: Ana Gretel Echazú Böschemeier
Published: (2023)
Similar Items
-
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025) -
Interleaved Head Attention
by: Duvvuri, Sai Surya, et al.
Published: (2026) -
Repeat After Me: Transformers are Better than State Space Models at Copying
by: Jelassi, Samy, et al.
Published: (2024) -
Universal Length Generalization with Turing Programs
by: Hou, Kaiying, et al.
Published: (2024) -
The Role of Sparsity for Length Generalization in Transformers
by: Golowich, Noah, et al.
Published: (2025)