Learning When to Attend: Conditional Memory Access for Long-Context LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Choudhary, Sakshi, Chattopadhyay, Aditya, Zancato, Luca, Nunez, Elvis, Trager, Matthew, Xia, Wei, Soatto, Stefano |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
por: Peng, Liangzu, et al.
Publicado: (2025)
por: Peng, Liangzu, et al.
Publicado: (2025)
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
por: Nunez, Elvis, et al.
Publicado: (2024)
por: Nunez, Elvis, et al.
Publicado: (2024)
PICASO: Permutation-Invariant Context Composition with State Space Models
por: Liu, Tian Yu, et al.
Publicado: (2025)
por: Liu, Tian Yu, et al.
Publicado: (2025)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
por: Zancato, Luca, et al.
Publicado: (2024)
por: Zancato, Luca, et al.
Publicado: (2024)
Multi-Modal Hallucination Control by Visual Information Grounding
por: Favero, Alessandro, et al.
Publicado: (2024)
por: Favero, Alessandro, et al.
Publicado: (2024)
Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
por: Trager, Matthew, et al.
Publicado: (2023)
por: Trager, Matthew, et al.
Publicado: (2023)
Compositional Structures in Neural Embedding and Interaction Decompositions
por: Trager, Matthew, et al.
Publicado: (2024)
por: Trager, Matthew, et al.
Publicado: (2024)
Maximally-Informative Retrieval for State Space Model Generation
por: Becker, Evan, et al.
Publicado: (2025)
por: Becker, Evan, et al.
Publicado: (2025)
Priming: Hybrid State Space Models From Pre-trained Transformers
por: Chattopadhyay, Aditya, et al.
Publicado: (2026)
por: Chattopadhyay, Aditya, et al.
Publicado: (2026)
Experience-Guided Adaptation of Inference-Time Reasoning Strategies
por: Stein, Adam, et al.
Publicado: (2025)
por: Stein, Adam, et al.
Publicado: (2025)
Learning to Focus: Focal Attention for Selective and Scalable Transformers
por: Ram, Dhananjay, et al.
Publicado: (2025)
por: Ram, Dhananjay, et al.
Publicado: (2025)
Descriminative-Generative Custom Tokens for Vision-Language Models
por: Perera, Pramuditha, et al.
Publicado: (2025)
por: Perera, Pramuditha, et al.
Publicado: (2025)
e1: Learning Adaptive Control of Reasoning Effort
por: Kleinman, Michael, et al.
Publicado: (2025)
por: Kleinman, Michael, et al.
Publicado: (2025)
LATTS: Locally Adaptive Test-Time Scaling
por: Uscidda, Theo, et al.
Publicado: (2025)
por: Uscidda, Theo, et al.
Publicado: (2025)
The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
por: Stewart, Lawrence, et al.
Publicado: (2024)
por: Stewart, Lawrence, et al.
Publicado: (2024)
Cycles of Thought: Measuring LLM Confidence through Stable Explanations
por: Becker, Evan, et al.
Publicado: (2024)
por: Becker, Evan, et al.
Publicado: (2024)
Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models
por: Liu, Xiaoze, et al.
Publicado: (2026)
por: Liu, Xiaoze, et al.
Publicado: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
por: Xiao, Chaojun, et al.
Publicado: (2024)
por: Xiao, Chaojun, et al.
Publicado: (2024)
EvoMAS: Evolutionary Generation of Multi-Agent Systems
por: Hu, Yuntong, et al.
Publicado: (2026)
por: Hu, Yuntong, et al.
Publicado: (2026)
OASIS: Online Activation Subspace Learning for Memory-Efficient Training
por: Choudhary, Sakshi, et al.
Publicado: (2026)
por: Choudhary, Sakshi, et al.
Publicado: (2026)
When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal
por: Phalod, Aditya Ajay
Publicado: (2026)
por: Phalod, Aditya Ajay
Publicado: (2026)
Eigen Attention: Attention in Low-Rank Space for KV Cache Compression
por: Saxena, Utkarsh, et al.
Publicado: (2024)
por: Saxena, Utkarsh, et al.
Publicado: (2024)
Long-Short Alignment for Effective Long-Context Modeling in LLMs
por: Du, Tianqi, et al.
Publicado: (2025)
por: Du, Tianqi, et al.
Publicado: (2025)
Learning When Not to Attend Globally
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
por: Nagaraj, Manish, et al.
Publicado: (2025)
por: Nagaraj, Manish, et al.
Publicado: (2025)
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly
por: Hosseini, Peyman, et al.
Publicado: (2024)
por: Hosseini, Peyman, et al.
Publicado: (2024)
Attendre: Wait To Attend By Retrieval With Evicted Queries in Memory-Based Transformers for Long Context Processing
por: Yang, Zi, et al.
Publicado: (2024)
por: Yang, Zi, et al.
Publicado: (2024)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
por: Zhang, Qingru, et al.
Publicado: (2023)
por: Zhang, Qingru, et al.
Publicado: (2023)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
por: Kim, Minsoo, et al.
Publicado: (2024)
por: Kim, Minsoo, et al.
Publicado: (2024)
LongSafety: Enhance Safety for Long-Context LLMs
por: Huang, Mianqiu, et al.
Publicado: (2024)
por: Huang, Mianqiu, et al.
Publicado: (2024)
How LLMs Might Think
por: Gottlieb, Joseph, et al.
Publicado: (2026)
por: Gottlieb, Joseph, et al.
Publicado: (2026)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
por: Li, Zeju, et al.
Publicado: (2026)
por: Li, Zeju, et al.
Publicado: (2026)
Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs
por: Bansal, Rachit, et al.
Publicado: (2025)
por: Bansal, Rachit, et al.
Publicado: (2025)
A Controlled Study on Long Context Extension and Generalization in LLMs
por: Lu, Yi, et al.
Publicado: (2024)
por: Lu, Yi, et al.
Publicado: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
por: Shang, Xianpeng, et al.
Publicado: (2026)
por: Shang, Xianpeng, et al.
Publicado: (2026)
Toward Conversational Agents with Context and Time Sensitive Long-term Memory
por: Alonso, Nick, et al.
Publicado: (2024)
por: Alonso, Nick, et al.
Publicado: (2024)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
por: He, Zifan, et al.
Publicado: (2024)
por: He, Zifan, et al.
Publicado: (2024)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
por: Sivtsov, Danil, et al.
Publicado: (2025)
por: Sivtsov, Danil, et al.
Publicado: (2025)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
por: Jiang, Huiqiang, et al.
Publicado: (2023)
por: Jiang, Huiqiang, et al.
Publicado: (2023)
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
por: Bai, Yushi, et al.
Publicado: (2024)
por: Bai, Yushi, et al.
Publicado: (2024)
Ejemplares similares
-
Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
por: Peng, Liangzu, et al.
Publicado: (2025) -
Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models
por: Nunez, Elvis, et al.
Publicado: (2024) -
PICASO: Permutation-Invariant Context Composition with State Space Models
por: Liu, Tian Yu, et al.
Publicado: (2025) -
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
por: Zancato, Luca, et al.
Publicado: (2024) -
Multi-Modal Hallucination Control by Visual Information Grounding
por: Favero, Alessandro, et al.
Publicado: (2024)