Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
Fuente:
arXiv
Guardado en:
| Autores principales: | Omidi, Parsa, Huang, Xingshuai, Laborieux, Axel, Nikpour, Bahareh, Shi, Tianyu, Eshaghi, Armaghan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
por: Huang, Xingshuai, et al.
Publicado: (2026)
por: Huang, Xingshuai, et al.
Publicado: (2026)
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
por: Wang, Xindi, et al.
Publicado: (2024)
por: Wang, Xindi, et al.
Publicado: (2024)
Language-Guided Reinforcement Learning for Hard Attention in Few-Shot Learning
por: Nikpour, Bahareh, et al.
Publicado: (2023)
por: Nikpour, Bahareh, et al.
Publicado: (2023)
Improving equilibrium propagation without weight symmetry through Jacobian homeostasis
por: Laborieux, Axel, et al.
Publicado: (2023)
por: Laborieux, Axel, et al.
Publicado: (2023)
Leveraging Distillation Techniques for Document Understanding: A Case Study with FLAN-T5
por: Lamott, Marcel, et al.
Publicado: (2024)
por: Lamott, Marcel, et al.
Publicado: (2024)
Theories of synaptic memory consolidation and intelligent plasticity for continual learning
por: Zenke, Friedemann, et al.
Publicado: (2024)
por: Zenke, Friedemann, et al.
Publicado: (2024)
An Evolved Universal Transformer Memory
por: Cetin, Edoardo, et al.
Publicado: (2024)
por: Cetin, Edoardo, et al.
Publicado: (2024)
Enhancing Length Extrapolation in Sequential Models with Pointer-Augmented Neural Memory
por: Le, Hung, et al.
Publicado: (2024)
por: Le, Hung, et al.
Publicado: (2024)
Review, Remask, Refine (R3): Process-Guided Block Diffusion for Text Generation
por: Mounier, Nikita, et al.
Publicado: (2025)
por: Mounier, Nikita, et al.
Publicado: (2025)
Selective Attention: Enhancing Transformer through Principled Context Control
por: Zhang, Xuechen, et al.
Publicado: (2024)
por: Zhang, Xuechen, et al.
Publicado: (2024)
Goal-Conditioned Data Augmentation for Offline Reinforcement Learning
por: Huang, Xingshuai, et al.
Publicado: (2024)
por: Huang, Xingshuai, et al.
Publicado: (2024)
Design Principle Transfer in Neural Architecture Search via Large Language Models
por: Zhou, Xun, et al.
Publicado: (2024)
por: Zhou, Xun, et al.
Publicado: (2024)
LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization
por: Li, Junsong, et al.
Publicado: (2025)
por: Li, Junsong, et al.
Publicado: (2025)
Towards Principled Design of Mixture-of-Experts Language Models under Memory and Inference Constraints
por: Liew, Seng Pei, et al.
Publicado: (2026)
por: Liew, Seng Pei, et al.
Publicado: (2026)
DRDT3: Diffusion-Refined Decision Test-Time Training Model
por: Huang, Xingshuai, et al.
Publicado: (2025)
por: Huang, Xingshuai, et al.
Publicado: (2025)
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers
por: Neshaei, Seyed Parsa, et al.
Publicado: (2024)
por: Neshaei, Seyed Parsa, et al.
Publicado: (2024)
PanGu-$π$: Enhancing Language Model Architectures via Nonlinearity Compensation
por: Wang, Yunhe, et al.
Publicado: (2023)
por: Wang, Yunhe, et al.
Publicado: (2023)
Efficacy of Large Language Models in Systematic Reviews
por: Shah, Aaditya, et al.
Publicado: (2024)
por: Shah, Aaditya, et al.
Publicado: (2024)
A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
por: Oche, Agada Joseph, et al.
Publicado: (2025)
por: Oche, Agada Joseph, et al.
Publicado: (2025)
Can Memory-Augmented Language Models Generalize on Reasoning-in-a-Haystack Tasks?
por: Das, Payel, et al.
Publicado: (2025)
por: Das, Payel, et al.
Publicado: (2025)
Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling
por: Acharya, Rishiraj
Publicado: (2025)
por: Acharya, Rishiraj
Publicado: (2025)
Decoupling Scores and Text: The Politeness Principle in Peer Review
por: Wen, Yingxuan
Publicado: (2026)
por: Wen, Yingxuan
Publicado: (2026)
On the Power of Convolution Augmented Transformer
por: Li, Mingchen, et al.
Publicado: (2024)
por: Li, Mingchen, et al.
Publicado: (2024)
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
por: Ha, Hyeonjeong, et al.
Publicado: (2026)
por: Ha, Hyeonjeong, et al.
Publicado: (2026)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
por: Huang, Zeyi, et al.
Publicado: (2026)
por: Huang, Zeyi, et al.
Publicado: (2026)
Towards Robust Few-Shot Text Classification Using Transformer Architectures and Dual Loss Strategies
por: Han, Xu, et al.
Publicado: (2025)
por: Han, Xu, et al.
Publicado: (2025)
Data Augmentations for Improved (Large) Language Model Generalization
por: Feder, Amir, et al.
Publicado: (2023)
por: Feder, Amir, et al.
Publicado: (2023)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
por: Jobanputra, Mayank, et al.
Publicado: (2025)
por: Jobanputra, Mayank, et al.
Publicado: (2025)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
por: Vendrell, Victor Conchello, et al.
Publicado: (2026)
por: Vendrell, Victor Conchello, et al.
Publicado: (2026)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
por: Liu, Haokun, et al.
Publicado: (2025)
por: Liu, Haokun, et al.
Publicado: (2025)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
por: Huang, Yixiao, et al.
Publicado: (2025)
por: Huang, Yixiao, et al.
Publicado: (2025)
Learning Novel Transformer Architecture for Time-series Forecasting
por: Zhang, Juyuan, et al.
Publicado: (2025)
por: Zhang, Juyuan, et al.
Publicado: (2025)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
por: Qiu, Zeju, et al.
Publicado: (2026)
por: Qiu, Zeju, et al.
Publicado: (2026)
DTRNet: Dynamic Token Routing Network to Reduce Quadratic Costs in Transformers
por: Sharma, Aman, et al.
Publicado: (2025)
por: Sharma, Aman, et al.
Publicado: (2025)
Linking In-context Learning in Transformers to Human Episodic Memory
por: Ji-An, Li, et al.
Publicado: (2024)
por: Ji-An, Li, et al.
Publicado: (2024)
Efficient Context Propagating Perceiver Architectures for Auto-Regressive Language Modeling
por: Mahmood, Kaleel, et al.
Publicado: (2024)
por: Mahmood, Kaleel, et al.
Publicado: (2024)
Chain-of-Defensive-Thought: Structured Reasoning Elicits Robustness in Large Language Models against Reference Corruption
por: Wang, Wenxiao, et al.
Publicado: (2025)
por: Wang, Wenxiao, et al.
Publicado: (2025)
Modeling Bilingual Sentence Processing: Evaluating RNN and Transformer Architectures for Cross-Language Structural Priming
por: Zhang, Demi, et al.
Publicado: (2024)
por: Zhang, Demi, et al.
Publicado: (2024)
A Systematic Review of Federated Generative Models
por: Gargary, Ashkan Vedadi, et al.
Publicado: (2024)
por: Gargary, Ashkan Vedadi, et al.
Publicado: (2024)
MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
por: Xia, Kejing, et al.
Publicado: (2026)
por: Xia, Kejing, et al.
Publicado: (2026)
Ejemplares similares
-
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
por: Huang, Xingshuai, et al.
Publicado: (2026) -
Beyond the Limits: A Survey of Techniques to Extend the Context Length in Large Language Models
por: Wang, Xindi, et al.
Publicado: (2024) -
Language-Guided Reinforcement Learning for Hard Attention in Few-Shot Learning
por: Nikpour, Bahareh, et al.
Publicado: (2023) -
Improving equilibrium propagation without weight symmetry through Jacobian homeostasis
por: Laborieux, Axel, et al.
Publicado: (2023) -
Leveraging Distillation Techniques for Document Understanding: A Case Study with FLAN-T5
por: Lamott, Marcel, et al.
Publicado: (2024)