Attendre: Wait To Attend By Retrieval With Evicted Queries in Memory-Based Transformers for Long Context Processing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Zi, Hua, Nan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Equipping Transformer with Random-Access Reading for Long-Context Understanding
von: Yang, Chenghao, et al.
Veröffentlicht: (2024)
von: Yang, Chenghao, et al.
Veröffentlicht: (2024)
Query-focused and Memory-aware Reranker for Long Context Processing
von: Li, Yuqing, et al.
Veröffentlicht: (2026)
von: Li, Yuqing, et al.
Veröffentlicht: (2026)
Retrieval Or Holistic Understanding? Dolce: Differentiate Our Long Context Evaluation Tasks
von: Yang, Zi
Veröffentlicht: (2024)
von: Yang, Zi
Veröffentlicht: (2024)
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
von: Choudhary, Sakshi, et al.
Veröffentlicht: (2026)
von: Choudhary, Sakshi, et al.
Veröffentlicht: (2026)
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
von: Li, Jinze, et al.
Veröffentlicht: (2026)
von: Li, Jinze, et al.
Veröffentlicht: (2026)
Learning to Evict from Key-Value Cache
von: Moschella, Luca, et al.
Veröffentlicht: (2026)
von: Moschella, Luca, et al.
Veröffentlicht: (2026)
HMT: Hierarchical Memory Transformer for Efficient Long Context Language Processing
von: He, Zifan, et al.
Veröffentlicht: (2024)
von: He, Zifan, et al.
Veröffentlicht: (2024)
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
von: Qian, Hongjin, et al.
Veröffentlicht: (2024)
von: Qian, Hongjin, et al.
Veröffentlicht: (2024)
Query-Focused Retrieval Heads Improve Long-Context Reasoning and Re-ranking
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
von: Zhang, Wuwei, et al.
Veröffentlicht: (2025)
LongEmbed: Extending Embedding Models for Long Context Retrieval
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
von: Zhu, Dawei, et al.
Veröffentlicht: (2024)
EMS: Adaptive Evict-then-Merge Strategy for Head-wise KV Cache Compression Based on Global-Local Importance
von: Li, Yingxin, et al.
Veröffentlicht: (2024)
von: Li, Yingxin, et al.
Veröffentlicht: (2024)
Infinite Retrieval: Attention Enhanced LLMs in Long-Context Processing
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
von: Ye, Xiaoju, et al.
Veröffentlicht: (2025)
OkraLong: A Flexible Retrieval-Augmented Framework for Long-Text Query Processing
von: Hui, Yulong, et al.
Veröffentlicht: (2025)
von: Hui, Yulong, et al.
Veröffentlicht: (2025)
WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems
von: Yu, Jiangnan, et al.
Veröffentlicht: (2026)
von: Yu, Jiangnan, et al.
Veröffentlicht: (2026)
HTAM: Hierarchical Transition-Attended Memory for Operator Optimization
von: Zhang, Yining, et al.
Veröffentlicht: (2026)
von: Zhang, Yining, et al.
Veröffentlicht: (2026)
HingeMem: Boundary Guided Long-Term Memory with Query Adaptive Retrieval for Scalable Dialogues
von: Zhong, Yijie, et al.
Veröffentlicht: (2026)
von: Zhong, Yijie, et al.
Veröffentlicht: (2026)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
Query Suggestion for Retrieval-Augmented Generation via Dynamic In-Context Learning
von: Spaeh, Fabian, et al.
Veröffentlicht: (2026)
von: Spaeh, Fabian, et al.
Veröffentlicht: (2026)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
von: Song, Woomin, et al.
Veröffentlicht: (2025)
von: Song, Woomin, et al.
Veröffentlicht: (2025)
Learning to Retrieve In-Context Examples for Large Language Models
von: Wang, Liang, et al.
Veröffentlicht: (2023)
von: Wang, Liang, et al.
Veröffentlicht: (2023)
S$^3$-Attention:Attention-Aligned Endogenous Retrieval for Memory-Bounded Long-Context Inference
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
von: Ma, Qingsen, et al.
Veröffentlicht: (2026)
Evaluating Long-Term Memory for Long-Context Question Answering
von: Terranova, Alessandra, et al.
Veröffentlicht: (2025)
von: Terranova, Alessandra, et al.
Veröffentlicht: (2025)
Improving Retrieval in Sponsored Search by Leveraging Query Context Signals
von: Mohankumar, Akash Kumar, et al.
Veröffentlicht: (2024)
von: Mohankumar, Akash Kumar, et al.
Veröffentlicht: (2024)
QUITO: Accelerating Long-Context Reasoning through Query-Guided Context Compression
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
von: Wang, Wenshan, et al.
Veröffentlicht: (2024)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
von: Shlapentokh-Rothman, Michal, et al.
Veröffentlicht: (2026)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
von: Liu, Di, et al.
Veröffentlicht: (2024)
von: Liu, Di, et al.
Veröffentlicht: (2024)
Retrieval Head Mechanistically Explains Long-Context Factuality
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
von: Wu, Wenhao, et al.
Veröffentlicht: (2024)
Inference Scaling for Long-Context Retrieval Augmented Generation
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
von: Yue, Zhenrui, et al.
Veröffentlicht: (2024)
Evaluating Multilingual Long-Context Models for Retrieval and Reasoning
von: Agrawal, Ameeta, et al.
Veröffentlicht: (2024)
von: Agrawal, Ameeta, et al.
Veröffentlicht: (2024)
Unstructured Evidence Attribution for Long Context Query Focused Summarization
von: Wright, Dustin, et al.
Veröffentlicht: (2025)
von: Wright, Dustin, et al.
Veröffentlicht: (2025)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
von: Liu, Weijie, et al.
Veröffentlicht: (2024)
von: Liu, Weijie, et al.
Veröffentlicht: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
von: Xiao, Guangxuan, et al.
Veröffentlicht: (2024)
XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference
von: Monteiro, João, et al.
Veröffentlicht: (2024)
von: Monteiro, João, et al.
Veröffentlicht: (2024)
Gated Differentiable Working Memory for Long-Context Language Modeling
von: Mei, Lingrui, et al.
Veröffentlicht: (2026)
von: Mei, Lingrui, et al.
Veröffentlicht: (2026)
Literary Evidence Retrieval via Long-Context Language Models
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
von: Thai, Katherine, et al.
Veröffentlicht: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
von: Chen, Guanzheng, et al.
Veröffentlicht: (2025)
Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agents
von: Pollertlam, Natchanon, et al.
Veröffentlicht: (2026)
von: Pollertlam, Natchanon, et al.
Veröffentlicht: (2026)
Chunk, Align, Select: A Simple Long-sequence Processing Method for Transformers
von: Xie, Jiawen, et al.
Veröffentlicht: (2023)
von: Xie, Jiawen, et al.
Veröffentlicht: (2023)
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices
von: Chen, Jiyu, et al.
Veröffentlicht: (2025)
von: Chen, Jiyu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Equipping Transformer with Random-Access Reading for Long-Context Understanding
von: Yang, Chenghao, et al.
Veröffentlicht: (2024) -
Query-focused and Memory-aware Reranker for Long Context Processing
von: Li, Yuqing, et al.
Veröffentlicht: (2026) -
Retrieval Or Holistic Understanding? Dolce: Differentiate Our Long Context Evaluation Tasks
von: Yang, Zi
Veröffentlicht: (2024) -
Learning When to Attend: Conditional Memory Access for Long-Context LLMs
von: Choudhary, Sakshi, et al.
Veröffentlicht: (2026) -
OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
von: Li, Jinze, et al.
Veröffentlicht: (2026)