ATLAS: Learning to Optimally Memorize the Context at Test Time
Fuente:
arXiv
Guardado en:
| Autores principales: | Behrouz, Ali, Li, Zeman, Kacham, Praneeth, Daliri, Majid, Deng, Yuan, Zhong, Peilin, Razaviyayn, Meisam, Mirrokni, Vahab |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TNT: Improving Chunkwise Training for Test-Time Memorization
por: Li, Zeman, et al.
Publicado: (2025)
por: Li, Zeman, et al.
Publicado: (2025)
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
por: Behrouz, Ali, et al.
Publicado: (2025)
por: Behrouz, Ali, et al.
Publicado: (2025)
Titans: Learning to Memorize at Test Time
por: Behrouz, Ali, et al.
Publicado: (2024)
por: Behrouz, Ali, et al.
Publicado: (2024)
Nested Learning: The Illusion of Deep Learning Architectures
por: Behrouz, Ali, et al.
Publicado: (2025)
por: Behrouz, Ali, et al.
Publicado: (2025)
Memory Caching: RNNs with Growing Memory
por: Behrouz, Ali, et al.
Publicado: (2026)
por: Behrouz, Ali, et al.
Publicado: (2026)
Trellis: Learning to Compress Key-Value Memory in Attention Models
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
por: Li, Zeman, et al.
Publicado: (2024)
por: Li, Zeman, et al.
Publicado: (2024)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
por: Li, Zeman, et al.
Publicado: (2025)
por: Li, Zeman, et al.
Publicado: (2025)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
por: Kacham, Praneeth, et al.
Publicado: (2023)
por: Kacham, Praneeth, et al.
Publicado: (2023)
Sampling and Loss Weights in Multi-Domain Training
por: Salmani, Mahdi, et al.
Publicado: (2025)
por: Salmani, Mahdi, et al.
Publicado: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
por: Han, Insu, et al.
Publicado: (2025)
por: Han, Insu, et al.
Publicado: (2025)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
por: Zandieh, Amir, et al.
Publicado: (2025)
por: Zandieh, Amir, et al.
Publicado: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
por: Nguyen, Dang, et al.
Publicado: (2025)
por: Nguyen, Dang, et al.
Publicado: (2025)
Optimal Communication for Classic Functions in the Coordinator Model and Beyond
por: Esfandiari, Hossein, et al.
Publicado: (2024)
por: Esfandiari, Hossein, et al.
Publicado: (2024)
High-Dimensional Geometric Streaming for Nearly Low Rank Data
por: Esfandiari, Hossein, et al.
Publicado: (2024)
por: Esfandiari, Hossein, et al.
Publicado: (2024)
Understanding the Role of Training Data in Test-Time Scaling
por: Javanmard, Adel, et al.
Publicado: (2025)
por: Javanmard, Adel, et al.
Publicado: (2025)
ECO: Quantized Training without Full-Precision Master Weights
por: Nikdan, Mahdi, et al.
Publicado: (2026)
por: Nikdan, Mahdi, et al.
Publicado: (2026)
Early Stopping for Large Reasoning Models via Confidence Dynamics
por: Hosseini, Parsa, et al.
Publicado: (2026)
por: Hosseini, Parsa, et al.
Publicado: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
por: Das, Rudrajit, et al.
Publicado: (2026)
por: Das, Rudrajit, et al.
Publicado: (2026)
Memorization in In-Context Learning
por: Golchin, Shahriar, et al.
Publicado: (2024)
por: Golchin, Shahriar, et al.
Publicado: (2024)
Unlocking the Theory Behind Scaling 1-Bit Neural Networks
por: Daliri, Majid, et al.
Publicado: (2024)
por: Daliri, Majid, et al.
Publicado: (2024)
Lattice: Learning to Efficiently Compress the Memory
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Context Memorization for Efficient Long Context Generation
por: Okoshi, Yasuyuki, et al.
Publicado: (2026)
por: Okoshi, Yasuyuki, et al.
Publicado: (2026)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
por: Karami, Mahdi, et al.
Publicado: (2025)
por: Karami, Mahdi, et al.
Publicado: (2025)
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
por: Niu, Jingcheng, et al.
Publicado: (2025)
por: Niu, Jingcheng, et al.
Publicado: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
por: Javanmard, Adel, et al.
Publicado: (2026)
por: Javanmard, Adel, et al.
Publicado: (2026)
Optimistic Rates for Learning from Label Proportions
por: Li, Gene, et al.
Publicado: (2024)
por: Li, Gene, et al.
Publicado: (2024)
Generalization or Memorization? Brittleness Testing for Chess-Trained Language Models
por: Tang, Ethan
Publicado: (2026)
por: Tang, Ethan
Publicado: (2026)
SubGen: Token Generation in Sublinear Time and Memory
por: Zandieh, Amir, et al.
Publicado: (2024)
por: Zandieh, Amir, et al.
Publicado: (2024)
Rezwan: Leveraging Large Language Models for Comprehensive Hadith Text Processing: A 1.2M Corpus Development
por: Asgari-Bidhendi, Majid, et al.
Publicado: (2025)
por: Asgari-Bidhendi, Majid, et al.
Publicado: (2025)
Undesirable Memorization in Large Language Models: A Survey
por: Satvaty, Ali, et al.
Publicado: (2024)
por: Satvaty, Ali, et al.
Publicado: (2024)
Optimal Differentially Private Model Training with Public Data
por: Lowy, Andrew, et al.
Publicado: (2023)
por: Lowy, Andrew, et al.
Publicado: (2023)
Towards Better Generalization in Open-Domain Question Answering by Mitigating Context Memorization
por: Zhang, Zixuan, et al.
Publicado: (2024)
por: Zhang, Zixuan, et al.
Publicado: (2024)
When to Memorize and When to Stop: Gated Recurrent Memory for Long-Context Reasoning
por: Sheng, Leheng, et al.
Publicado: (2026)
por: Sheng, Leheng, et al.
Publicado: (2026)
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction
por: Vadlapati, Praneeth
Publicado: (2024)
por: Vadlapati, Praneeth
Publicado: (2024)
A Dual-Directional Context-Aware Test-Time Learning for Text Classification
por: Xu, Dong, et al.
Publicado: (2025)
por: Xu, Dong, et al.
Publicado: (2025)
FastMem: Fast Memorization of Prompt Improves Context Awareness of Large Language Models
por: Zhu, Junyi, et al.
Publicado: (2024)
por: Zhu, Junyi, et al.
Publicado: (2024)
AutoPureData: Automated Filtering of Undesirable Web Data to Update LLM Knowledge
por: Vadlapati, Praneeth
Publicado: (2024)
por: Vadlapati, Praneeth
Publicado: (2024)
Memorization or Interpolation ? Detecting LLM Memorization through Input Perturbation Analysis
por: Djiré, Albérick Euraste, et al.
Publicado: (2025)
por: Djiré, Albérick Euraste, et al.
Publicado: (2025)
Ejemplares similares
-
TNT: Improving Chunkwise Training for Test-Time Memorization
por: Li, Zeman, et al.
Publicado: (2025) -
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
por: Behrouz, Ali, et al.
Publicado: (2025) -
Titans: Learning to Memorize at Test Time
por: Behrouz, Ali, et al.
Publicado: (2024) -
Nested Learning: The Illusion of Deep Learning Architectures
por: Behrouz, Ali, et al.
Publicado: (2025) -
Memory Caching: RNNs with Growing Memory
por: Behrouz, Ali, et al.
Publicado: (2026)