FlashBack:Efficient Retrieval-Augmented Language Modeling for Long Context Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Runheng, Xiao, Xingchen, Huang, Heyan, Chi, Zewen, Wu, Zhijing |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026)
by: Liu, Runheng, et al.
Published: (2026)
Training-free Truthfulness Detection via Value Vectors in LLMs
by: Liu, Runheng, et al.
Published: (2025)
by: Liu, Runheng, et al.
Published: (2025)
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
by: Xiao, Xingchen, et al.
Published: (2026)
by: Xiao, Xingchen, et al.
Published: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
by: Zhou, Youchao, et al.
Published: (2024)
by: Zhou, Youchao, et al.
Published: (2024)
Inference Scaling for Long-Context Retrieval Augmented Generation
by: Yue, Zhenrui, et al.
Published: (2024)
by: Yue, Zhenrui, et al.
Published: (2024)
Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
by: Laitenberger, Alex, et al.
Published: (2025)
by: Laitenberger, Alex, et al.
Published: (2025)
RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding
by: Chen, Guanzheng, et al.
Published: (2025)
by: Chen, Guanzheng, et al.
Published: (2025)
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models
by: Su, Weihang, et al.
Published: (2024)
by: Su, Weihang, et al.
Published: (2024)
BGE Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models
by: Luo, Kun, et al.
Published: (2024)
by: Luo, Kun, et al.
Published: (2024)
Multi-modal Retrieval Augmented Multi-modal Generation: Datasets, Evaluation Metrics and Strong Baselines
by: Ma, Zi-Ao, et al.
Published: (2024)
by: Ma, Zi-Ao, et al.
Published: (2024)
DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads
by: Xiao, Guangxuan, et al.
Published: (2024)
by: Xiao, Guangxuan, et al.
Published: (2024)
ZigzagAttention: Efficient Long-Context Inference with Exclusive Retrieval and Streaming Heads
by: Liu, Zhuorui, et al.
Published: (2025)
by: Liu, Zhuorui, et al.
Published: (2025)
Debate, Reflect, and Distill: Multi-Agent Feedback with Tree-Structured Preference Optimization for Efficient Language Model Enhancement
by: Zhou, Xiaofeng, et al.
Published: (2025)
by: Zhou, Xiaofeng, et al.
Published: (2025)
Exploring Fine-Tuning for In-Context Retrieval and Efficient KV-Caching in Long-Context Language Models
by: Molfese, Francesco Maria, et al.
Published: (2026)
by: Molfese, Francesco Maria, et al.
Published: (2026)
EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding Models
by: Xie, Jincheng, et al.
Published: (2026)
by: Xie, Jincheng, et al.
Published: (2026)
Deterministic Reversible Data Augmentation for Neural Machine Translation
by: Yao, Jiashu, et al.
Published: (2024)
by: Yao, Jiashu, et al.
Published: (2024)
Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
by: Hu, Xiang, et al.
Published: (2024)
by: Hu, Xiang, et al.
Published: (2024)
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024)
by: Huang, Lianming, et al.
Published: (2024)
Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection
by: Zhu, Yun, et al.
Published: (2024)
by: Zhu, Yun, et al.
Published: (2024)
FlashRAG: A Modular Toolkit for Efficient Retrieval-Augmented Generation Research
by: Jin, Jiajie, et al.
Published: (2024)
by: Jin, Jiajie, et al.
Published: (2024)
Mindscape-Aware Retrieval Augmented Generation for Improved Long Context Understanding
by: Li, Yuqing, et al.
Published: (2025)
by: Li, Yuqing, et al.
Published: (2025)
D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
by: Wan, Zhongwei, et al.
Published: (2024)
by: Wan, Zhongwei, et al.
Published: (2024)
Membership Inference Attack against Long-Context Large Language Models
by: Wang, Zixiong, et al.
Published: (2024)
by: Wang, Zixiong, et al.
Published: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
Lighter And Better: Towards Flexible Context Adaptation For Retrieval Augmented Generation
by: Liu, Zheng, et al.
Published: (2024)
by: Liu, Zheng, et al.
Published: (2024)
A MapReduce Approach to Effectively Utilize Long Context Information in Retrieval Augmented Language Models
by: Zhang, Gongbo, et al.
Published: (2024)
by: Zhang, Gongbo, et al.
Published: (2024)
UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Retrieval Head Mechanistically Explains Long-Context Factuality
by: Wu, Wenhao, et al.
Published: (2024)
by: Wu, Wenhao, et al.
Published: (2024)
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
FlashDecoding++: Faster Large Language Model Inference on GPUs
by: Hong, Ke, et al.
Published: (2023)
by: Hong, Ke, et al.
Published: (2023)
FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
by: Hu, Zhanqiu, et al.
Published: (2025)
by: Hu, Zhanqiu, et al.
Published: (2025)
Literary Evidence Retrieval via Long-Context Language Models
by: Thai, Katherine, et al.
Published: (2025)
by: Thai, Katherine, et al.
Published: (2025)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
by: Liu, Di, et al.
Published: (2024)
by: Liu, Di, et al.
Published: (2024)
Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs
by: Yu, Shuyang, et al.
Published: (2024)
by: Yu, Shuyang, et al.
Published: (2024)
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation
by: Qian, Hongjin, et al.
Published: (2024)
by: Qian, Hongjin, et al.
Published: (2024)
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation
by: Qiu, Yuli, et al.
Published: (2024)
by: Qiu, Yuli, et al.
Published: (2024)
Grounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation
by: Chen, Jiamin, et al.
Published: (2025)
by: Chen, Jiamin, et al.
Published: (2025)
Retrieval meets Long Context Large Language Models
by: Xu, Peng, et al.
Published: (2023)
by: Xu, Peng, et al.
Published: (2023)
Enhancing Robustness of Retrieval-Augmented Language Models with In-Context Learning
by: Park, Seong-Il, et al.
Published: (2024)
by: Park, Seong-Il, et al.
Published: (2024)
Similar Items
-
Zero-Shot Detection of LLM-Generated Text via Implicit Reward Model
by: Liu, Runheng, et al.
Published: (2026) -
Training-free Truthfulness Detection via Value Vectors in LLMs
by: Liu, Runheng, et al.
Published: (2025) -
MASS-RAG: Multi-Agent Synthesis Retrieval-Augmented Generation
by: Xiao, Xingchen, et al.
Published: (2026) -
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025) -
Subtopic-aware View Sampling and Temporal Aggregation for Long-form Document Matching
by: Zhou, Youchao, et al.
Published: (2024)