MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Jaiswal, Ajay, Hannah, Lauren, Kim, Han-Byul, Hoang, Duc, Kundu, Arnav, Farajtabar, Mehrdad, Cho, Minsik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
di: Kim, Minsoo, et al.
Pubblicazione: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
di: Hoang, Duc, et al.
Pubblicazione: (2026)
di: Hoang, Duc, et al.
Pubblicazione: (2026)
M+: Extending MemoryLLM with Scalable Long-Term Memory
di: Wang, Yu, et al.
Pubblicazione: (2025)
di: Wang, Yu, et al.
Pubblicazione: (2025)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
di: Alizadeh, Keivan, et al.
Pubblicazione: (2023)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2023)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
R2 Loss: Range Restriction Loss for Model Compression and Quantization
di: Kundu, Arnav, et al.
Pubblicazione: (2023)
di: Kundu, Arnav, et al.
Pubblicazione: (2023)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
di: Lyu, Sisuo, et al.
Pubblicazione: (2026)
di: Lyu, Sisuo, et al.
Pubblicazione: (2026)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
di: Cao, Jiaqi, et al.
Pubblicazione: (2025)
di: Cao, Jiaqi, et al.
Pubblicazione: (2025)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2024)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
di: Alizadeh, Keivan, et al.
Pubblicazione: (2024)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
di: Sarkar, Aditya, et al.
Pubblicazione: (2026)
di: Sarkar, Aditya, et al.
Pubblicazione: (2026)
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
di: Qu, Yuwen, et al.
Pubblicazione: (2026)
di: Qu, Yuwen, et al.
Pubblicazione: (2026)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
di: Nishu, Kumari, et al.
Pubblicazione: (2025)
di: Nishu, Kumari, et al.
Pubblicazione: (2025)
Self-supervised Deep Hyperspectral Inpainting with the Plug and Play and Deep Image Prior Models
di: Li, Shuo, et al.
Pubblicazione: (2025)
di: Li, Shuo, et al.
Pubblicazione: (2025)
Analysis and Synthesis Denoisers for Forward-Backward Plug-and-Play Algorithms
di: Kowalski, Matthieu, et al.
Pubblicazione: (2024)
di: Kowalski, Matthieu, et al.
Pubblicazione: (2024)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
di: Yang, Ke, et al.
Pubblicazione: (2026)
di: Yang, Ke, et al.
Pubblicazione: (2026)
Romanization-Induced Mispronunciations in Korean: How Latin Letters Alter the Perception of Japanese Voiceless Consonants
di: Kang, Byul
Pubblicazione: (2025)
di: Kang, Byul
Pubblicazione: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
di: Cho, Minsik, et al.
Pubblicazione: (2024)
di: Cho, Minsik, et al.
Pubblicazione: (2024)
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
di: Kim, HyunJin, et al.
Pubblicazione: (2023)
di: Kim, HyunJin, et al.
Pubblicazione: (2023)
MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
di: Huang, Yizhe, et al.
Pubblicazione: (2025)
di: Huang, Yizhe, et al.
Pubblicazione: (2025)
Uniform boundedness on rational maps with automorphisms
di: Han, Minsik
Pubblicazione: (2024)
di: Han, Minsik
Pubblicazione: (2024)
A Study of Student Dependency on Artificial Intelligence Applications in their Education: With Reference to Indore City
di: Ajay Jaiswal
Pubblicazione: (2025)
di: Ajay Jaiswal
Pubblicazione: (2025)
Online Temporal Action Localization with Memory-Augmented Transformer
di: Song, Youngkil, et al.
Pubblicazione: (2024)
di: Song, Youngkil, et al.
Pubblicazione: (2024)
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm
di: Hou, Zhichao, et al.
Pubblicazione: (2024)
di: Hou, Zhichao, et al.
Pubblicazione: (2024)
Safe Memory Reclamation Techniques
di: Singh, Ajay
Pubblicazione: (2025)
di: Singh, Ajay
Pubblicazione: (2025)
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
di: Kim, Injae, et al.
Pubblicazione: (2026)
di: Kim, Injae, et al.
Pubblicazione: (2026)
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
di: Sarawgi, Utkarsh Oggy, et al.
Pubblicazione: (2023)
di: Sarawgi, Utkarsh Oggy, et al.
Pubblicazione: (2023)
Plug-and-Play Transformer Modules for Test-Time Adaptation
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
di: Chang, Xiangyu, et al.
Pubblicazione: (2024)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
di: Jeong, Yoonwoo, et al.
Pubblicazione: (2023)
di: Jeong, Yoonwoo, et al.
Pubblicazione: (2023)
Plug-n-Play Three Pulse Twin Field QKD
di: Gayathri, Anagha, et al.
Pubblicazione: (2025)
di: Gayathri, Anagha, et al.
Pubblicazione: (2025)
Merging Feed-Forward Sublayers for Compressed Transformers
di: Verma, Neha, et al.
Pubblicazione: (2025)
di: Verma, Neha, et al.
Pubblicazione: (2025)
Documenti analoghi
-
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026) -
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
di: Kim, Han-Byul, et al.
Pubblicazione: (2025) -
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
di: Kim, Minsoo, et al.
Pubblicazione: (2025) -
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025) -
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)