ATLAS: Learning to Optimally Memorize the Context at Test Time
Fuente:
arXiv
Saved in:
| Main Authors: | Behrouz, Ali, Li, Zeman, Kacham, Praneeth, Daliri, Majid, Deng, Yuan, Zhong, Peilin, Razaviyayn, Meisam, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024)
by: Behrouz, Ali, et al.
Published: (2024)
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025)
by: Behrouz, Ali, et al.
Published: (2025)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
by: Li, Zeman, et al.
Published: (2024)
by: Li, Zeman, et al.
Published: (2024)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
by: Li, Zeman, et al.
Published: (2025)
by: Li, Zeman, et al.
Published: (2025)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
by: Kacham, Praneeth, et al.
Published: (2023)
by: Kacham, Praneeth, et al.
Published: (2023)
Sampling and Loss Weights in Multi-Domain Training
by: Salmani, Mahdi, et al.
Published: (2025)
by: Salmani, Mahdi, et al.
Published: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
by: Zandieh, Amir, et al.
Published: (2025)
by: Zandieh, Amir, et al.
Published: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
by: Nguyen, Dang, et al.
Published: (2025)
by: Nguyen, Dang, et al.
Published: (2025)
Optimal Communication for Classic Functions in the Coordinator Model and Beyond
by: Esfandiari, Hossein, et al.
Published: (2024)
by: Esfandiari, Hossein, et al.
Published: (2024)
High-Dimensional Geometric Streaming for Nearly Low Rank Data
by: Esfandiari, Hossein, et al.
Published: (2024)
by: Esfandiari, Hossein, et al.
Published: (2024)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
ECO: Quantized Training without Full-Precision Master Weights
by: Nikdan, Mahdi, et al.
Published: (2026)
by: Nikdan, Mahdi, et al.
Published: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
by: Hosseini, Parsa, et al.
Published: (2026)
by: Hosseini, Parsa, et al.
Published: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
by: Das, Rudrajit, et al.
Published: (2026)
by: Das, Rudrajit, et al.
Published: (2026)
Memorization in In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2024)
by: Golchin, Shahriar, et al.
Published: (2024)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
by: Daliri, Majid, et al.
Published: (2024)
by: Daliri, Majid, et al.
Published: (2024)
Lattice: Learning to Efficiently Compress the Memory
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Context Memorization for Efficient Long Context Generation
by: Okoshi, Yasuyuki, et al.
Published: (2026)
by: Okoshi, Yasuyuki, et al.
Published: (2026)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
by: Karami, Mahdi, et al.
Published: (2025)
by: Karami, Mahdi, et al.
Published: (2025)
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
by: Niu, Jingcheng, et al.
Published: (2025)
by: Niu, Jingcheng, et al.
Published: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Optimistic Rates for Learning from Label Proportions
by: Li, Gene, et al.
Published: (2024)
by: Li, Gene, et al.
Published: (2024)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
by: Tang, Ethan
Published: (2026)
by: Tang, Ethan
Published: (2026)
SubGen: Token Generation in Sublinear Time and Memory
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Rezwan: Leveraging Large Language Models for Comprehensive Hadith Text Processing: A 1.2M Corpus Development
by: Asgari-Bidhendi, Majid, et al.
Published: (2025)
by: Asgari-Bidhendi, Majid, et al.
Published: (2025)
Undesirable Memorization in Large Language Models: A Survey
by: Satvaty, Ali, et al.
Published: (2024)
by: Satvaty, Ali, et al.
Published: (2024)
Optimal Differentially Private Model Training with Public Data
by: Lowy, Andrew, et al.
Published: (2023)
by: Lowy, Andrew, et al.
Published: (2023)
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization
by: Zhang, Zixuan, et al.
Published: (2024)
by: Zhang, Zixuan, et al.
Published: (2024)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
by: Sheng, Leheng, et al.
Published: (2026)
by: Sheng, Leheng, et al.
Published: (2026)
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
by: Vadlapati, Praneeth
Published: (2024)
by: Vadlapati, Praneeth
Published: (2024)
A Dual-Directional Context-Aware Test-Time Learning for Text Classification
by: Xu, Dong, et al.
Published: (2025)
by: Xu, Dong, et al.
Published: (2025)
FastMem: Fast Memorization of Prompt Improves Context Awareness of Large Language Models
by: Zhu, Junyi, et al.
Published: (2024)
by: Zhu, Junyi, et al.
Published: (2024)
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
by: Vadlapati, Praneeth
Published: (2024)
by: Vadlapati, Praneeth
Published: (2024)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
by: Djiré, Albérick Euraste, et al.
Published: (2025)
by: Djiré, Albérick Euraste, et al.
Published: (2025)
Similar Items
-
TNT: Improving Chunkwise Training for Test-Time Memorization
by: Li, Zeman, et al.
Published: (2025) -
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
by: Behrouz, Ali, et al.
Published: (2025) -
Titans: Learning to Memorize at Test Time
by: Behrouz, Ali, et al.
Published: (2024) -
Nested Learning: The Illusion of Deep Learning Architectures
by: Behrouz, Ali, et al.
Published: (2025) -
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)