RASD: Retrieval-Augmented Speculative Decoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Quan, Guofeng, Feng, Wenfeng, Hao, Chuzhan, Jiang, Guochao, Zhang, Yuewei, Wang, Hao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search
von: Hao, Chuzhan, et al.
Veröffentlicht: (2026)
von: Hao, Chuzhan, et al.
Veröffentlicht: (2026)
Mixture-of-LoRAs: An Efficient Multitask Tuning for Large Language Models
von: Feng, Wenfeng, et al.
Veröffentlicht: (2024)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2024)
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
von: Hao, Chuzhan, et al.
Veröffentlicht: (2025)
von: Hao, Chuzhan, et al.
Veröffentlicht: (2025)
DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
von: Jiang, Guochao, et al.
Veröffentlicht: (2026)
von: Jiang, Guochao, et al.
Veröffentlicht: (2026)
FlowKV: A Disaggregated Inference Framework with Low-Latency KV Cache Transfer and Load-Aware Scheduling
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
von: Li, Weiqing, et al.
Veröffentlicht: (2025)
PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
von: Zhang, Zihong, et al.
Veröffentlicht: (2026)
FlashThink: An Early Exit Method For Efficient Reasoning
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
von: Jiang, Guochao, et al.
Veröffentlicht: (2025)
Online Speculative Decoding
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2023)
When, What, and How: Rethinking Retrieval-Enhanced Speculative Decoding
von: Fang, Min, et al.
Veröffentlicht: (2025)
von: Fang, Min, et al.
Veröffentlicht: (2025)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Speculative RAG: Enhancing Retrieval Augmented Generation through Drafting
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
von: Wang, Zilong, et al.
Veröffentlicht: (2024)
Speculative Decoding for Multi-Sample Inference
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
von: Li, Yiwei, et al.
Veröffentlicht: (2025)
Unraveling and Mitigating Retriever Inconsistencies in Retrieval-Augmented Large Language Models
von: Li, Mingda, et al.
Veröffentlicht: (2024)
von: Li, Mingda, et al.
Veröffentlicht: (2024)
CREST: Effectively Compacting a Datastore For Retrieval-Based Speculative Decoding
von: Ho, Sophia, et al.
Veröffentlicht: (2024)
von: Ho, Sophia, et al.
Veröffentlicht: (2024)
Cross-Attention Speculative Decoding
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
von: Zhong, Wei, et al.
Veröffentlicht: (2025)
Batch Speculative Decoding Done Right
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
von: Zhang, Ranran Haoran, et al.
Veröffentlicht: (2025)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
von: Brown, Oscar, et al.
Veröffentlicht: (2024)
REST: Retrieval-Based Speculative Decoding
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
von: He, Zhenyu, et al.
Veröffentlicht: (2023)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Speculative Decoding: Performance or Illusion?
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoxuan, et al.
Veröffentlicht: (2025)
Mamba Drafters for Speculative Decoding
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
von: Choi, Daewon, et al.
Veröffentlicht: (2025)
Constrained Decoding with Speculative Lookaheads
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
von: Nakshatri, Nishanth, et al.
Veröffentlicht: (2024)
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
von: Jin, Tao, et al.
Veröffentlicht: (2026)
von: Jin, Tao, et al.
Veröffentlicht: (2026)
PACER: Blockwise Pre-verification for Speculative Decoding with Adaptive Length
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
von: Zhang, Situo, et al.
Veröffentlicht: (2026)
Cacheback: Speculative Decoding With Nothing But Cache
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
von: Ma, Zhiyao, et al.
Veröffentlicht: (2025)
Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering
von: Wang, Changjian, et al.
Veröffentlicht: (2025)
von: Wang, Changjian, et al.
Veröffentlicht: (2025)
SAM Decoding: Speculative Decoding via Suffix Automaton
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Hu, Yuxuan, et al.
Veröffentlicht: (2024)
SENSE: Semantic Embedding Navigation with Soft-gated Evaluation for Retrieval-based Speculative Decoding
von: Chen, Shaowen, et al.
Veröffentlicht: (2026)
von: Chen, Shaowen, et al.
Veröffentlicht: (2026)
Hybrid Verified Decoding: Learning to Allocate Verification in Speculative Decoding
von: Su, Xin, et al.
Veröffentlicht: (2026)
von: Su, Xin, et al.
Veröffentlicht: (2026)
Mixture of Attentions For Speculative Decoding
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
von: Zimmer, Matthieu, et al.
Veröffentlicht: (2024)
Traversal Verification for Speculative Tree Decoding
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
Parallel Context-of-Experts Decoding for Retrieval Augmented Generation
von: Corallo, Giulio, et al.
Veröffentlicht: (2026)
von: Corallo, Giulio, et al.
Veröffentlicht: (2026)
Entropy-Aware Speculative Decoding Toward Improved LLM Reasoning
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
von: Su, Tiancheng, et al.
Veröffentlicht: (2025)
Evaluation of Retrieval-Augmented Generation: A Survey
von: Yu, Hao, et al.
Veröffentlicht: (2024)
von: Yu, Hao, et al.
Veröffentlicht: (2024)
Speculate, then Collaborate: Fusing Knowledge of Language Models during Decoding
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
RAPID: Efficient Retrieval-Augmented Long Text Generation with Writing Planning and Information Discovery
von: Gu, Hongchao, et al.
Veröffentlicht: (2025)
von: Gu, Hongchao, et al.
Veröffentlicht: (2025)
Accelerating Large Language Model Reasoning via Speculative Search
von: Wang, Zhihai, et al.
Veröffentlicht: (2025)
von: Wang, Zhihai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AirRAG: Autonomous Strategic Planning and Reasoning Steer Retrieval Augmented Generation
von: Feng, Wenfeng, et al.
Veröffentlicht: (2025) -
Beyond Stochastic Exploration: What Makes Training Data Valuable for Agentic Search
von: Hao, Chuzhan, et al.
Veröffentlicht: (2026) -
Mixture-of-LoRAs: An Efficient Multitask Tuning for Large Language Models
von: Feng, Wenfeng, et al.
Veröffentlicht: (2024) -
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models
von: Jiang, Guochao, et al.
Veröffentlicht: (2025) -
DynaSearcher: Dynamic Knowledge Graph Augmented Search Agent via Multi-Reward Reinforcement Learning
von: Hao, Chuzhan, et al.
Veröffentlicht: (2025)