MHA-RAG: Improving Efficiency, Accuracy, and Consistency by Encoding Exemplars as Soft Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Jain, Abhinav, Yao, Xinyu, Reps, Thomas, Jermaine, Christopher |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
by: Jain, Abhinav, et al.
Published: (2024)
by: Jain, Abhinav, et al.
Published: (2024)
RAG-Modulo: Solving Sequential Tasks using Experience, Critics, and Language Models
by: Jain, Abhinav, et al.
Published: (2024)
by: Jain, Abhinav, et al.
Published: (2024)
On The Finetuning of MLIPs Through the Lens of Iterated Maps With BPTT
by: Dramko, Evan, et al.
Published: (2025)
by: Dramko, Evan, et al.
Published: (2025)
A-MHA*: Anytime Multi-Heuristic A*
by: Natarajan, Ramkumar, et al.
Published: (2025)
by: Natarajan, Ramkumar, et al.
Published: (2025)
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
by: Zheng, Longtao, et al.
Published: (2023)
by: Zheng, Longtao, et al.
Published: (2023)
KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs
by: Sanyal, Debopam, et al.
Published: (2026)
by: Sanyal, Debopam, et al.
Published: (2026)
Beyond Self-Consistency: Ensemble Reasoning Boosts Consistency and Accuracy of LLMs in Cancer Staging
by: Chang, Chia-Hsuan, et al.
Published: (2024)
by: Chang, Chia-Hsuan, et al.
Published: (2024)
Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars
by: Wu, Zhaoxuan, et al.
Published: (2024)
by: Wu, Zhaoxuan, et al.
Published: (2024)
Mixture of Attention Schemes (MoAS): Learning to Route Between MHA, GQA, and MQA
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Lateral Tree-of-Thoughts Surpasses ToT by Incorporating Logically-Consistent, Low-Utility Candidates
by: Madahar, Abhinav
Published: (2025)
by: Madahar, Abhinav
Published: (2025)
COM-BOM: Bayesian Exemplar Search for Efficiently Exploring the Accuracy-Calibration Pareto Frontier
by: Luo, Gaoxiang, et al.
Published: (2025)
by: Luo, Gaoxiang, et al.
Published: (2025)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers
by: Sawarkar, Kunal, et al.
Published: (2024)
by: Sawarkar, Kunal, et al.
Published: (2024)
Soft Self-Consistency Improves Language Model Agents
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
Automated Tensor-Relational Decomposition for Large-Scale Sparse Tensor Computation
by: Tang, Yuxin, et al.
Published: (2026)
by: Tang, Yuxin, et al.
Published: (2026)
Contradiction Detection in RAG Systems: Evaluating LLMs as Context Validators for Improved Information Consistency
by: Gokul, Vignesh, et al.
Published: (2025)
by: Gokul, Vignesh, et al.
Published: (2025)
Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion
by: Zhang, Sen, et al.
Published: (2026)
by: Zhang, Sen, et al.
Published: (2026)
Improving Statistical Significance in Human Evaluation of Automatic Metrics via Soft Pairwise Accuracy
by: Thompson, Brian, et al.
Published: (2024)
by: Thompson, Brian, et al.
Published: (2024)
E^2GraphRAG: Streamlining Graph-based RAG for High Efficiency and Effectiveness
by: Zhao, Yibo, et al.
Published: (2025)
by: Zhao, Yibo, et al.
Published: (2025)
DANA: Domain-Aware Neurosymbolic Agents for Consistency and Accuracy
by: Luong, Vinh, et al.
Published: (2024)
by: Luong, Vinh, et al.
Published: (2024)
Compressed Causal Reasoning: Quantization and GraphRAG Effects on Interventional and Counterfactual Accuracy
by: Nwaiwu, Steve, et al.
Published: (2025)
by: Nwaiwu, Steve, et al.
Published: (2025)
Dual Prompt-Driven Feature Encoding for Nighttime UAV Tracking
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025)
by: Chawla, Kushal, et al.
Published: (2025)
Aligning LLMs for the Classroom with Knowledge-Based Retrieval -- A Comparative RAG Study
by: Jain, Amay, et al.
Published: (2025)
by: Jain, Amay, et al.
Published: (2025)
Learning to Rotate: Temporal and Semantic Rotary Encoding for Sequential Modeling
by: Cheng, Hailing, et al.
Published: (2026)
by: Cheng, Hailing, et al.
Published: (2026)
Consistent and Distinctive: LLM Benchmark Efficiency via Maximum Independent Set Prompt Selection on Similarity Graphs
by: Kjorvezir, Denica, et al.
Published: (2026)
by: Kjorvezir, Denica, et al.
Published: (2026)
Enhancing Retrieval and Managing Retrieval: A Four-Module Synergy for Improved Quality and Efficiency in RAG Systems
by: Shi, Yunxiao, et al.
Published: (2024)
by: Shi, Yunxiao, et al.
Published: (2024)
PEA: Enhancing LLM Performance on Computational-Reasoning Tasks
by: Wang, Zi, et al.
Published: (2025)
by: Wang, Zi, et al.
Published: (2025)
Enhancing RAG Efficiency with Adaptive Context Compression
by: Guo, Shuyu, et al.
Published: (2025)
by: Guo, Shuyu, et al.
Published: (2025)
Learning to Love Edge Cases in Formative Math Assessment: Using the AMMORE Dataset and Chain-of-Thought Prompting to Improve Grading Accuracy
by: Henkel, Owen, et al.
Published: (2024)
by: Henkel, Owen, et al.
Published: (2024)
Exemplar-condensed Federated Class-incremental Learning
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
From RAG to RICHES: Retrieval Interlaced with Sequence Generation
by: Jain, Palak, et al.
Published: (2024)
by: Jain, Palak, et al.
Published: (2024)
Enhancing LLM-Based Text Classification in Political Science: Automatic Prompt Optimization and Dynamic Exemplar Selection for Few-Shot Learning
by: Liu, Menglin, et al.
Published: (2024)
by: Liu, Menglin, et al.
Published: (2024)
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
by: Cui, Christopher Z., et al.
Published: (2026)
by: Cui, Christopher Z., et al.
Published: (2026)
RaFe: Ranking Feedback Improves Query Rewriting for RAG
by: Mao, Shengyu, et al.
Published: (2024)
by: Mao, Shengyu, et al.
Published: (2024)
Experience as a Compass: Multi-agent RAG with Evolving Orchestration and Agent Prompts
by: Li, Sha, et al.
Published: (2026)
by: Li, Sha, et al.
Published: (2026)
Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist
by: Kamijo, Tatsuya, et al.
Published: (2026)
by: Kamijo, Tatsuya, et al.
Published: (2026)
Characterizing Motion Encoding in Video Diffusion Timesteps
by: Baherwani, Vatsal, et al.
Published: (2025)
by: Baherwani, Vatsal, et al.
Published: (2025)
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Similar Items
-
Prompt Tuning Strikes Back: Customizing Foundation Models with Low-Rank Prompt Adaptation
by: Jain, Abhinav, et al.
Published: (2024) -
RAG-Modulo: Solving Sequential Tasks using Experience, Critics, and Language Models
by: Jain, Abhinav, et al.
Published: (2024) -
On The Finetuning of MLIPs Through the Lens of Iterated Maps With BPTT
by: Dramko, Evan, et al.
Published: (2025) -
A-MHA*: Anytime Multi-Heuristic A*
by: Natarajan, Ramkumar, et al.
Published: (2025) -
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control
by: Zheng, Longtao, et al.
Published: (2023)