Understanding Transformer from the Perspective of Associative Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Shu, Xu, Mingyu, Ao, Tenglong, Shi, Guang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
by: Swain, Kabir, et al.
Published: (2026)
by: Swain, Kabir, et al.
Published: (2026)
Understanding Representation of Deep Equilibrium Models from Neural Collapse Perspective
by: Sun, Haixiang, et al.
Published: (2024)
by: Sun, Haixiang, et al.
Published: (2024)
MeSH: Memory-as-State-Highways for Recursive Transformers
by: Yu, Chengting, et al.
Published: (2025)
by: Yu, Chengting, et al.
Published: (2025)
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023)
by: Deiseroth, Björn, et al.
Published: (2023)
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
by: Wang, Weixin, et al.
Published: (2025)
by: Wang, Weixin, et al.
Published: (2025)
The Memory Perturbation Equation: Understanding Model's Sensitivity to Data
by: Nickl, Peter, et al.
Published: (2023)
by: Nickl, Peter, et al.
Published: (2023)
Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation
by: Yang, Yuchen, et al.
Published: (2024)
by: Yang, Yuchen, et al.
Published: (2024)
Hybrid Associative Memories
by: Lufkin, Leon, et al.
Published: (2026)
by: Lufkin, Leon, et al.
Published: (2026)
Adaptive Advantage-Guided Policy Regularization for Offline Reinforcement Learning
by: Liu, Tenglong, et al.
Published: (2024)
by: Liu, Tenglong, et al.
Published: (2024)
Recurrent Action Transformer with Memory
by: Cherepanov, Egor, et al.
Published: (2023)
by: Cherepanov, Egor, et al.
Published: (2023)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Associative Recurrent Memory Transformer
by: Rodkin, Ivan, et al.
Published: (2024)
by: Rodkin, Ivan, et al.
Published: (2024)
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
by: Wu, Xiefeng, et al.
Published: (2025)
by: Wu, Xiefeng, et al.
Published: (2025)
Learning Associative Memories with Gradient Descent
by: Cabannes, Vivien, et al.
Published: (2024)
by: Cabannes, Vivien, et al.
Published: (2024)
Memory Caching: RNNs with Growing Memory
by: Behrouz, Ali, et al.
Published: (2026)
by: Behrouz, Ali, et al.
Published: (2026)
Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning Tasks
by: Xu, Xingcheng, et al.
Published: (2024)
by: Xu, Xingcheng, et al.
Published: (2024)
Transformer-Based Spatial-Temporal Counterfactual Outcomes Estimation
by: Li, He, et al.
Published: (2025)
by: Li, He, et al.
Published: (2025)
Parallel BiLSTM-Transformer networks for forecasting chaotic dynamics
by: Ma, Junwen, et al.
Published: (2025)
by: Ma, Junwen, et al.
Published: (2025)
Honest Lying: Understanding Memory Confabulation in Reflexive Agents
by: Dixit, Prakhar, et al.
Published: (2026)
by: Dixit, Prakhar, et al.
Published: (2026)
Diffusion Policies with Value-Conditional Optimization for Offline Reinforcement Learning
by: Ma, Yunchang, et al.
Published: (2025)
by: Ma, Yunchang, et al.
Published: (2025)
Skill Expansion and Composition in Parameter Space
by: Liu, Tenglong, et al.
Published: (2025)
by: Liu, Tenglong, et al.
Published: (2025)
Context Distillation as Latent Memory Management
by: Zheng, Ziyang, et al.
Published: (2026)
by: Zheng, Ziyang, et al.
Published: (2026)
LENS: Large Pre-trained Transformer for Exploring Financial Time Series Regularities
by: Xu, Yuanjian, et al.
Published: (2024)
by: Xu, Yuanjian, et al.
Published: (2024)
It Ain't That Bad: Understanding the Mysterious Performance Drop in OOD Generalization for Generative Transformer Models
by: Xu, Xingcheng, et al.
Published: (2023)
by: Xu, Xingcheng, et al.
Published: (2023)
Re:Frame -- Retrieving Experience From Associative Memory
by: Zelezetsky, Daniil, et al.
Published: (2025)
by: Zelezetsky, Daniil, et al.
Published: (2025)
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents
by: Liu, Jiaqi, et al.
Published: (2026)
by: Liu, Jiaqi, et al.
Published: (2026)
Echo State Transformer: Attention Over Finite Memories
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
by: Zhang, Yukun, et al.
Published: (2024)
by: Zhang, Yukun, et al.
Published: (2024)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
by: Ke, Yekun, et al.
Published: (2024)
by: Ke, Yekun, et al.
Published: (2024)
Understanding and Mitigating Overrefusal in LLMs from an Unveiling Perspective of Safety Decision Boundary
by: Pan, Licheng, et al.
Published: (2025)
by: Pan, Licheng, et al.
Published: (2025)
ESSAM: A Novel Competitive Evolution Strategies Approach to Reinforcement Learning for Memory Efficient LLMs Fine-Tuning
by: Sun, Zhishen, et al.
Published: (2026)
by: Sun, Zhishen, et al.
Published: (2026)
The Trap of Trajectory: Towards Understanding and Mitigating Spurious Correlations in Agentic Memory
by: Tang, Luoxi, et al.
Published: (2026)
by: Tang, Luoxi, et al.
Published: (2026)
Memory- and Latency-Constrained Inference of Large Language Models via Adaptive Split Computing
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization
by: Tan, Chengli, et al.
Published: (2025)
by: Tan, Chengli, et al.
Published: (2025)
A Flat Minima Perspective on Understanding Augmentations and Model Robustness
by: Yoo, Weebum, et al.
Published: (2025)
by: Yoo, Weebum, et al.
Published: (2025)
GPU Memory Requirement Prediction for Deep Learning Task Based on Bidirectional Gated Recurrent Unit Optimization Transformer
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
LOOKAT: Lookup-Optimized Key-Attention for Memory-Efficient Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
QFlash: Bridging Quantization and Memory Efficiency in Vision Transformer Attention
by: Oh, Sehyeon, et al.
Published: (2026)
by: Oh, Sehyeon, et al.
Published: (2026)
Retrieval-Augmented Decision Transformer: External Memory for In-context RL
by: Schmied, Thomas, et al.
Published: (2024)
by: Schmied, Thomas, et al.
Published: (2024)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
by: Qiu, Zeju, et al.
Published: (2026)
by: Qiu, Zeju, et al.
Published: (2026)
Similar Items
-
Tensor Cache: Eviction-conditioned Associative Memory for Transformers
by: Swain, Kabir, et al.
Published: (2026) -
Understanding Representation of Deep Equilibrium Models from Neural Collapse Perspective
by: Sun, Haixiang, et al.
Published: (2024) -
MeSH: Memory-as-State-Highways for Recursive Transformers
by: Yu, Chengting, et al.
Published: (2025) -
AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation
by: Deiseroth, Björn, et al.
Published: (2023) -
Rethinking Langevin Thompson Sampling from A Stochastic Approximation Perspective
by: Wang, Weixin, et al.
Published: (2025)