Accelerating Retrieval-Augmented Language Model Serving with Speculation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhihao, Zhu, Alan, Yang, Lijie, Xu, Yihua, Li, Lanting, Phothilimthana, Phitchaya Mangpo, Jia, Zhihao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Model Augmented Exercise Retrieval for Personalized Language Learning
by: Xu, Austin, et al.
Published: (2024)
by: Xu, Austin, et al.
Published: (2024)
REST: Retrieval-Based Speculative Decoding
by: He, Zhenyu, et al.
Published: (2023)
by: He, Zhenyu, et al.
Published: (2023)
NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research
by: Chandrasekhar, Achuth, et al.
Published: (2025)
by: Chandrasekhar, Achuth, et al.
Published: (2025)
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation
by: Li, Mufei, et al.
Published: (2024)
by: Li, Mufei, et al.
Published: (2024)
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation
by: Zhang, Chenghao, et al.
Published: (2025)
by: Zhang, Chenghao, et al.
Published: (2025)
Retrieval-Augmented Generation with Graphs (GraphRAG)
by: Han, Haoyu, et al.
Published: (2024)
by: Han, Haoyu, et al.
Published: (2024)
Soft Prompt Tuning for Augmenting Dense Retrieval with Large Language Models
by: Peng, Zhiyuan, et al.
Published: (2023)
by: Peng, Zhiyuan, et al.
Published: (2023)
Graph Chain-of-Thought: Augmenting Large Language Models by Reasoning on Graphs
by: Jin, Bowen, et al.
Published: (2024)
by: Jin, Bowen, et al.
Published: (2024)
Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models
by: Weller, Orion, et al.
Published: (2024)
by: Weller, Orion, et al.
Published: (2024)
SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification
by: Miao, Xupeng, et al.
Published: (2023)
by: Miao, Xupeng, et al.
Published: (2023)
Optimizing Multi-Stage Language Models for Effective Text Retrieval
by: Trung, Quang Hoang, et al.
Published: (2024)
by: Trung, Quang Hoang, et al.
Published: (2024)
Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators
by: Su, Zhengyang, et al.
Published: (2026)
by: Su, Zhengyang, et al.
Published: (2026)
GraphRAFT: Retrieval Augmented Fine-Tuning for Knowledge Graphs on Graph Databases
by: Clemedtson, Alfred, et al.
Published: (2025)
by: Clemedtson, Alfred, et al.
Published: (2025)
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
by: Kahardipraja, Patrick, et al.
Published: (2025)
by: Kahardipraja, Patrick, et al.
Published: (2025)
Retrieval meets Long Context Large Language Models
by: Xu, Peng, et al.
Published: (2023)
by: Xu, Peng, et al.
Published: (2023)
Enhancing Question Answering Precision with Optimized Vector Retrieval and Instructions
by: Yang, Lixiao, et al.
Published: (2024)
by: Yang, Lixiao, et al.
Published: (2024)
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
by: Xi, Zhiheng, et al.
Published: (2025)
by: Xi, Zhiheng, et al.
Published: (2025)
Large Language Model Can Be a Foundation for Hidden Rationale-Based Retrieval
by: Ji, Luo, et al.
Published: (2024)
by: Ji, Luo, et al.
Published: (2024)
GraphER: An Efficient Graph-Based Enrichment and Reranking Method for Retrieval-Augmented Generation
by: Miao, Ruizhong, et al.
Published: (2026)
by: Miao, Ruizhong, et al.
Published: (2026)
Stochastic RAG: End-to-End Retrieval-Augmented Generation through Expected Utility Maximization
by: Zamani, Hamed, et al.
Published: (2024)
by: Zamani, Hamed, et al.
Published: (2024)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Fine-Tuning Large Language Models and Evaluating Retrieval Methods for Improved Question Answering on Building Codes
by: Aqib, Mohammad, et al.
Published: (2025)
by: Aqib, Mohammad, et al.
Published: (2025)
GORAG: Graph-based Online Retrieval Augmented Generation for Dynamic Few-shot Social Media Text Classification
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Retrieval Augmented Generation for Domain-specific Question Answering
by: Sharma, Sanat, et al.
Published: (2024)
by: Sharma, Sanat, et al.
Published: (2024)
ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging
by: Verma, Neha, et al.
Published: (2026)
by: Verma, Neha, et al.
Published: (2026)
ACER: Automatic Language Model Context Extension via Retrieval
by: Gao, Luyu, et al.
Published: (2024)
by: Gao, Luyu, et al.
Published: (2024)
RAMQA: A Unified Framework for Retrieval-Augmented Multi-Modal Question Answering
by: Bai, Yang, et al.
Published: (2025)
by: Bai, Yang, et al.
Published: (2025)
HaS: Accelerating RAG through Homology-Aware Speculative Retrieval
by: Peng, Peng, et al.
Published: (2026)
by: Peng, Peng, et al.
Published: (2026)
CoRAG: Collaborative Retrieval-Augmented Generation
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
SPARC-RAG: Adaptive Sequential-Parallel Scaling with Context Management for Retrieval-Augmented Generation
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
Make Large Language Model a Better Ranker
by: Chao, Wen-Shuo, et al.
Published: (2024)
by: Chao, Wen-Shuo, et al.
Published: (2024)
Beyond Sequential Reranking: Reranker-Guided Search Improves Reasoning Intensive Retrieval
by: Xu, Haike, et al.
Published: (2025)
by: Xu, Haike, et al.
Published: (2025)
Enhancing Q&A Text Retrieval with Ranking Models: Benchmarking, fine-tuning and deploying Rerankers for RAG
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
by: Moreira, Gabriel de Souza P., et al.
Published: (2024)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
by: Xue, Jiaqi, et al.
Published: (2024)
by: Xue, Jiaqi, et al.
Published: (2024)
SLMRec: Distilling Large Language Models into Small for Sequential Recommendation
by: Xu, Wujiang, et al.
Published: (2024)
by: Xu, Wujiang, et al.
Published: (2024)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
by: Jung, Dongwon, et al.
Published: (2024)
by: Jung, Dongwon, et al.
Published: (2024)
Optimization of Retrieval-Augmented Generation Context with Outlier Detection
by: Bulgakov, Vitaly
Published: (2024)
by: Bulgakov, Vitaly
Published: (2024)
Similar Items
-
Large Language Model Augmented Exercise Retrieval for Personalized Language Learning
by: Xu, Austin, et al.
Published: (2024) -
REST: Retrieval-Based Speculative Decoding
by: He, Zhenyu, et al.
Published: (2023) -
NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research
by: Chandrasekhar, Achuth, et al.
Published: (2025) -
Simple Is Effective: The Roles of Graphs and Large Language Models in Knowledge-Graph-Based Retrieval-Augmented Generation
by: Li, Mufei, et al.
Published: (2024) -
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
by: Xu, Ran, et al.
Published: (2024)