It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
Fuente:
arXiv
Salvato in:
| Autori principali: | Behrouz, Ali, Razaviyayn, Meisam, Zhong, Peilin, Mirrokni, Vahab |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Nested Learning: The Illusion of Deep Learning Architectures
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
Titans: Learning to Memorize at Test Time
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
TNT: Improving Chunkwise Training for Test-Time Memorization
di: Li, Zeman, et al.
Pubblicazione: (2025)
di: Li, Zeman, et al.
Pubblicazione: (2025)
Memory Caching: RNNs with Growing Memory
di: Behrouz, Ali, et al.
Pubblicazione: (2026)
di: Behrouz, Ali, et al.
Pubblicazione: (2026)
ATLAS: Learning to Optimally Memorize the Context at Test Time
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
di: Behrouz, Ali, et al.
Pubblicazione: (2025)
Sampling and Loss Weights in Multi-Domain Training
di: Salmani, Mahdi, et al.
Pubblicazione: (2025)
di: Salmani, Mahdi, et al.
Pubblicazione: (2025)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
di: Li, Zeman, et al.
Pubblicazione: (2025)
di: Li, Zeman, et al.
Pubblicazione: (2025)
Understanding the Role of Training Data in Test-Time Scaling
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
di: Javanmard, Adel, et al.
Pubblicazione: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
di: Li, Zeman, et al.
Pubblicazione: (2024)
di: Li, Zeman, et al.
Pubblicazione: (2024)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
di: Das, Rudrajit, et al.
Pubblicazione: (2026)
di: Das, Rudrajit, et al.
Pubblicazione: (2026)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
Lattice: Learning to Efficiently Compress the Memory
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
di: Javanmard, Adel, et al.
Pubblicazione: (2026)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
di: Zhang, Xinwei, et al.
Pubblicazione: (2024)
di: Zhang, Xinwei, et al.
Pubblicazione: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
di: Karami, Mahdi, et al.
Pubblicazione: (2025)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
di: Kacham, Praneeth, et al.
Pubblicazione: (2023)
di: Kacham, Praneeth, et al.
Pubblicazione: (2023)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
di: Zandieh, Amir, et al.
Pubblicazione: (2025)
di: Zandieh, Amir, et al.
Pubblicazione: (2025)
SubGen: Token Generation in Sublinear Time and Memory
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
Optimistic Rates for Learning from Label Proportions
di: Li, Gene, et al.
Pubblicazione: (2024)
di: Li, Gene, et al.
Pubblicazione: (2024)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
di: Nguyen, Dang, et al.
Pubblicazione: (2025)
Differentially Private Synthetic Data Release for Topics API Outputs
di: Dick, Travis, et al.
Pubblicazione: (2025)
di: Dick, Travis, et al.
Pubblicazione: (2025)
ECO: Quantized Training without Full-Precision Master Weights
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
di: Nikdan, Mahdi, et al.
Pubblicazione: (2026)
PolarQuant: Quantizing KV Caches with Polar Transformation
di: Han, Insu, et al.
Pubblicazione: (2025)
di: Han, Insu, et al.
Pubblicazione: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
Attention is All You Need Until You Need Retention
di: Yaslioglu, M. Murat
Pubblicazione: (2025)
di: Yaslioglu, M. Murat
Pubblicazione: (2025)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
di: Javanmard, Adel, et al.
Pubblicazione: (2024)
di: Javanmard, Adel, et al.
Pubblicazione: (2024)
Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
Output Perturbation for Differentially Private Convex Optimization: Faster and More General
di: Lowy, Andrew, et al.
Pubblicazione: (2021)
di: Lowy, Andrew, et al.
Pubblicazione: (2021)
Private Stochastic Optimization With Large Worst-Case Lipschitz Parameter
di: Lowy, Andrew, et al.
Pubblicazione: (2022)
di: Lowy, Andrew, et al.
Pubblicazione: (2022)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
di: Sanford, Clayton, et al.
Pubblicazione: (2024)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
di: Rodkin, Ivan, et al.
Pubblicazione: (2025)
Neural Network-Based Score Estimation in Diffusion Models: Optimization and Generalization
di: Han, Yinbin, et al.
Pubblicazione: (2024)
di: Han, Yinbin, et al.
Pubblicazione: (2024)
From Teacher to Student: Tracking Memorization Through Model Distillation
di: Singh, Simardeep
Pubblicazione: (2025)
di: Singh, Simardeep
Pubblicazione: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
di: Qu, Leigang, et al.
Pubblicazione: (2025)
di: Qu, Leigang, et al.
Pubblicazione: (2025)
Critical Windows of Complexity Control: When Transformers Decide to Reason or Memorize
di: Ali, Sarwan
Pubblicazione: (2026)
di: Ali, Sarwan
Pubblicazione: (2026)
Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
di: Meskin, Asal, et al.
Pubblicazione: (2025)
di: Meskin, Asal, et al.
Pubblicazione: (2025)
Membership and Memorization in LLM Knowledge Distillation
di: Zhang, Ziqi, et al.
Pubblicazione: (2025)
di: Zhang, Ziqi, et al.
Pubblicazione: (2025)
The Pitfalls of Memorization: When Memorization Hurts Generalization
di: Bayat, Reza, et al.
Pubblicazione: (2024)
di: Bayat, Reza, et al.
Pubblicazione: (2024)
Best of Both Worlds: Advantages of Hybrid Graph Sequence Models
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
di: Behrouz, Ali, et al.
Pubblicazione: (2024)
DeepCrossAttention: Supercharging Transformer Residual Connections
di: Heddes, Mike, et al.
Pubblicazione: (2025)
di: Heddes, Mike, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Nested Learning: The Illusion of Deep Learning Architectures
di: Behrouz, Ali, et al.
Pubblicazione: (2025) -
Titans: Learning to Memorize at Test Time
di: Behrouz, Ali, et al.
Pubblicazione: (2024) -
TNT: Improving Chunkwise Training for Test-Time Memorization
di: Li, Zeman, et al.
Pubblicazione: (2025) -
Memory Caching: RNNs with Growing Memory
di: Behrouz, Ali, et al.
Pubblicazione: (2026) -
ATLAS: Learning to Optimally Memorize the Context at Test Time
di: Behrouz, Ali, et al.
Pubblicazione: (2025)