Superposition Prompting: Improving and Accelerating Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Merth, Thomas, Fu, Qichen, Rastegari, Mohammad, Najibi, Mahyar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
by: Bhendawade, Nikhil, et al.
Published: (2024)
by: Bhendawade, Nikhil, et al.
Published: (2024)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
by: Mozaffari, Mohammad, et al.
Published: (2026)
by: Mozaffari, Mohammad, et al.
Published: (2026)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
by: Mehta, Sachin, et al.
Published: (2024)
by: Mehta, Sachin, et al.
Published: (2024)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
by: Bhendawade, Nikhil, et al.
Published: (2025)
by: Bhendawade, Nikhil, et al.
Published: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
by: Hoang, Duc N. M, et al.
Published: (2023)
by: Hoang, Duc N. M, et al.
Published: (2023)
Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
by: Wen, Yuxin, et al.
Published: (2024)
by: Wen, Yuxin, et al.
Published: (2024)
Learning Retrieval Augmentation for Personalized Dialogue Generation
by: Huang, Qiushi, et al.
Published: (2024)
by: Huang, Qiushi, et al.
Published: (2024)
SeedLM: Compressing LLM Weights into Seeds of Pseudo-Random Generators
by: Shafipour, Rasoul, et al.
Published: (2024)
by: Shafipour, Rasoul, et al.
Published: (2024)
Multiple Abstraction Level Retrieve Augment Generation
by: Zheng, Zheng, et al.
Published: (2025)
by: Zheng, Zheng, et al.
Published: (2025)
Enhancing Frame Detection with Retrieval Augmented Generation
by: Diallo, Papa Abdou Karim Karou, et al.
Published: (2025)
by: Diallo, Papa Abdou Karim Karou, et al.
Published: (2025)
Agentic Retrieval-Augmented Generation for Time Series Analysis
by: Ravuru, Chidaksh, et al.
Published: (2024)
by: Ravuru, Chidaksh, et al.
Published: (2024)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
by: Divekar, Abhishek, et al.
Published: (2024)
by: Divekar, Abhishek, et al.
Published: (2024)
Disco-RAG: Discourse-Aware Retrieval-Augmented Generation
by: Liu, Dongqi, et al.
Published: (2026)
by: Liu, Dongqi, et al.
Published: (2026)
Superposition Yields Robust Neural Scaling
by: Liu, Yizhou, et al.
Published: (2025)
by: Liu, Yizhou, et al.
Published: (2025)
A Graph-Retrieval-Augmented Generation Framework Enhances Decision-Making in the Circular Economy
by: Zhao, Yang, et al.
Published: (2025)
by: Zhao, Yang, et al.
Published: (2025)
Mix-of-Granularity: Optimize the Chunking Granularity for Retrieval-Augmented Generation
by: Zhong, Zijie, et al.
Published: (2024)
by: Zhong, Zijie, et al.
Published: (2024)
Less is More: Denoising Knowledge Graphs For Retrieval Augmented Generation
by: Zheng, Yilun, et al.
Published: (2025)
by: Zheng, Yilun, et al.
Published: (2025)
An Analysis of Hyper-Parameter Optimization Methods for Retrieval Augmented Generation
by: Orbach, Matan, et al.
Published: (2025)
by: Orbach, Matan, et al.
Published: (2025)
Improving Consistency in Retrieval-Augmented Systems with Group Similarity Rewards
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Ad Auctions for LLMs via Retrieval Augmented Generation
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
by: Hajiaghayi, MohammadTaghi, et al.
Published: (2024)
Investigating Retrieval-Augmented Generation in Quranic Studies: A Study of 13 Open-Source Large Language Models
by: Khalila, Zahra, et al.
Published: (2025)
by: Khalila, Zahra, et al.
Published: (2025)
RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models
by: Huang, Jie, et al.
Published: (2023)
by: Huang, Jie, et al.
Published: (2023)
Pointwise Mutual Information as a Performance Gauge for Retrieval-Augmented Generation
by: Liu, Tianyu, et al.
Published: (2024)
by: Liu, Tianyu, et al.
Published: (2024)
Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation
by: Qi, Jirui, et al.
Published: (2024)
by: Qi, Jirui, et al.
Published: (2024)
Say Less, Mean More: Leveraging Pragmatics in Retrieval-Augmented Generation
by: Riaz, Haris, et al.
Published: (2025)
by: Riaz, Haris, et al.
Published: (2025)
Optimizing Retrieval-Augmented Generation: Analysis of Hyperparameter Impact on Performance and Efficiency
by: Ammar, Adel, et al.
Published: (2025)
by: Ammar, Adel, et al.
Published: (2025)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
by: Yang, Shiping, et al.
Published: (2025)
by: Yang, Shiping, et al.
Published: (2025)
Soft Prompt Tuning for Augmenting Dense Retrieval with Large Language Models
by: Peng, Zhiyuan, et al.
Published: (2023)
by: Peng, Zhiyuan, et al.
Published: (2023)
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
Graph of Records: Boosting Retrieval Augmented Generation for Long-context Summarization with Graphs
by: Zhang, Haozhen, et al.
Published: (2024)
by: Zhang, Haozhen, et al.
Published: (2024)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
HiPRAG: Hierarchical Process Rewards for Efficient Agentic Retrieval Augmented Generation
by: Wu, Peilin, et al.
Published: (2025)
by: Wu, Peilin, et al.
Published: (2025)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
by: Prahlad, Deeksha, et al.
Published: (2025)
by: Prahlad, Deeksha, et al.
Published: (2025)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
by: Seo, Minju, et al.
Published: (2024)
by: Seo, Minju, et al.
Published: (2024)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
by: Alizadeh, Keivan, et al.
Published: (2023)
by: Alizadeh, Keivan, et al.
Published: (2023)
DynaGRAG | Exploring the Topology of Information for Advancing Language Understanding and Generation in Graph Retrieval-Augmented Generation
by: Thakrar, Karishma
Published: (2024)
by: Thakrar, Karishma
Published: (2024)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
by: Jadon, Aryan, et al.
Published: (2025)
by: Jadon, Aryan, et al.
Published: (2025)
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation
by: Zhang, Chenghao, et al.
Published: (2025)
by: Zhang, Chenghao, et al.
Published: (2025)
Similar Items
-
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024) -
Speculative Streaming: Fast LLM Inference without Auxiliary Models
by: Bhendawade, Nikhil, et al.
Published: (2024) -
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
by: Mozaffari, Mohammad, et al.
Published: (2026) -
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
by: Mehta, Sachin, et al.
Published: (2024) -
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
by: Mehta, Sachin, et al.
Published: (2024)