Memory Caching: RNNs with Growing Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Behrouz, Ali, Li, Zeman, Deng, Yuan, Zhong, Peilin, Razaviyayn, Meisam, Mirrokni, Vahab |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
Nested Learning: The Illusion of Deep Learning Architectures
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
TNT: Improving Chunkwise Training for Test-Time Memorization
von: Li, Zeman, et al.
Veröffentlicht: (2025)
von: Li, Zeman, et al.
Veröffentlicht: (2025)
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
von: Li, Zeman, et al.
Veröffentlicht: (2024)
von: Li, Zeman, et al.
Veröffentlicht: (2024)
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
von: Li, Zeman, et al.
Veröffentlicht: (2025)
von: Li, Zeman, et al.
Veröffentlicht: (2025)
Titans: Learning to Memorize at Test Time
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
ATLAS: Learning to Optimally Memorize the Context at Test Time
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
von: Behrouz, Ali, et al.
Veröffentlicht: (2025)
Sampling and Loss Weights in Multi-Domain Training
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
von: Salmani, Mahdi, et al.
Veröffentlicht: (2025)
Lattice: Learning to Efficiently Compress the Memory
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Less is More: Convergence Benefits of Fewer Data Weight Updates over Longer Horizon
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
von: Das, Rudrajit, et al.
Veröffentlicht: (2026)
SubGen: Token Generation in Sublinear Time and Memory
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
von: Zandieh, Amir, et al.
Veröffentlicht: (2024)
Trellis: Learning to Compress Key-Value Memory in Attention Models
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
PolarQuant: Quantizing KV Caches with Polar Transformation
von: Han, Insu, et al.
Veröffentlicht: (2025)
von: Han, Insu, et al.
Veröffentlicht: (2025)
MS-SSM: A Multi-Scale State Space Model for Efficient Sequence Modeling
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
von: Karami, Mahdi, et al.
Veröffentlicht: (2025)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Optimistic Rates for Learning from Label Proportions
von: Li, Gene, et al.
Veröffentlicht: (2024)
von: Li, Gene, et al.
Veröffentlicht: (2024)
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels
von: Kacham, Praneeth, et al.
Veröffentlicht: (2023)
von: Kacham, Praneeth, et al.
Veröffentlicht: (2023)
DiSK: Differentially Private Optimizer with Simplified Kalman Filter for Noise Reduction
von: Zhang, Xinwei, et al.
Veröffentlicht: (2024)
von: Zhang, Xinwei, et al.
Veröffentlicht: (2024)
Differentially Private Synthetic Data Release for Topics API Outputs
von: Dick, Travis, et al.
Veröffentlicht: (2025)
von: Dick, Travis, et al.
Veröffentlicht: (2025)
Hydra: Dual Exponentiated Memory for Multivariate Time Series Analysis
von: Meskin, Asal, et al.
Veröffentlicht: (2025)
von: Meskin, Asal, et al.
Veröffentlicht: (2025)
ECO: Quantized Training without Full-Precision Master Weights
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
von: Nikdan, Mahdi, et al.
Veröffentlicht: (2026)
Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
von: Hosseini, Parsa, et al.
Veröffentlicht: (2026)
Optimal Differentially Private Model Training with Public Data
von: Lowy, Andrew, et al.
Veröffentlicht: (2023)
von: Lowy, Andrew, et al.
Veröffentlicht: (2023)
Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
von: Bui, Ngoc, et al.
Veröffentlicht: (2025)
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
von: Zandieh, Amir, et al.
Veröffentlicht: (2025)
von: Zandieh, Amir, et al.
Veröffentlicht: (2025)
Your Code Agent Can Grow Alongside You with Structured Memory
von: Deng, Yi-Xuan, et al.
Veröffentlicht: (2026)
von: Deng, Yi-Xuan, et al.
Veröffentlicht: (2026)
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
von: Swain, Kabir, et al.
Veröffentlicht: (2026)
Learning from Aggregate responses: Instance Level versus Bag Level Loss Functions
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
von: Javanmard, Adel, et al.
Veröffentlicht: (2024)
IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
von: Mao, Yuzhen, et al.
Veröffentlicht: (2026)
von: Mao, Yuzhen, et al.
Veröffentlicht: (2026)
MemoryKT: An Integrative Memory-and-Forgetting Method for Knowledge Tracing
von: Lin, Mingrong, et al.
Veröffentlicht: (2025)
von: Lin, Mingrong, et al.
Veröffentlicht: (2025)
Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
von: Liu, Andy Zeyi, et al.
Veröffentlicht: (2026)
Context Distillation as Latent Memory Management
von: Zheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Zheng, Ziyang, et al.
Veröffentlicht: (2026)
Understanding Transformer Reasoning Capabilities via Graph Algorithms
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
Were RNNs All We Needed?
von: Feng, Leo, et al.
Veröffentlicht: (2024)
von: Feng, Leo, et al.
Veröffentlicht: (2024)
Federated Nested Learning: Collaborative Training of Self-Referential Memories for Test-Time Adaptation
von: Chen, Hong, et al.
Veröffentlicht: (2026)
von: Chen, Hong, et al.
Veröffentlicht: (2026)
Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
von: Behrouz, Ali, et al.
Veröffentlicht: (2024)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
von: Behrouz, Ali, et al.
Veröffentlicht: (2025) -
Nested Learning: The Illusion of Deep Learning Architectures
von: Behrouz, Ali, et al.
Veröffentlicht: (2025) -
TNT: Improving Chunkwise Training for Test-Time Memorization
von: Li, Zeman, et al.
Veröffentlicht: (2025) -
Addax: Utilizing Zeroth-Order Gradients to Improve Memory Efficiency and Performance of SGD for Fine-Tuning Language Models
von: Li, Zeman, et al.
Veröffentlicht: (2024) -
PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts
von: Li, Zeman, et al.
Veröffentlicht: (2025)