METIS: Fast Quality-Aware RAG Systems with Configuration Adaptation
Fuente:
arXiv
Saved in:
| Main Authors: | Ray, Siddhant, Pan, Rui, Gu, Zhuohan, Du, Kuntai, Feng, Shaoting, Ananthanarayanan, Ganesh, Netravali, Ravi, Jiang, Junchen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
by: Du, Kuntai, et al.
Published: (2023)
by: Du, Kuntai, et al.
Published: (2023)
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024)
by: Gu, Zhuohan, et al.
Published: (2024)
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
by: Yao, Jiayi, et al.
Published: (2024)
by: Yao, Jiayi, et al.
Published: (2024)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)
by: Feng, Shaoting, et al.
Published: (2025)
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
by: Li, Hanchen, et al.
Published: (2024)
by: Li, Hanchen, et al.
Published: (2024)
Kairos: A Scalable Serving System for Physical AI
by: Dai, Yinwei, et al.
Published: (2026)
by: Dai, Yinwei, et al.
Published: (2026)
A Methodology for Evaluating RAG Systems: A Case Study On Configuration Dependency Validation
by: Simon, Sebastian, et al.
Published: (2024)
by: Simon, Sebastian, et al.
Published: (2024)
CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
by: Liu, Yuhan, et al.
Published: (2023)
by: Liu, Yuhan, et al.
Published: (2023)
TC-RAG:Turing-Complete RAG's Case study on Medical LLM Systems
by: Jiang, Xinke, et al.
Published: (2024)
by: Jiang, Xinke, et al.
Published: (2024)
Practical RAG Evaluation: A Rarity-Aware Set-Based Metric and Cost-Latency-Quality Trade-offs
by: Dallaire, Etienne
Published: (2025)
by: Dallaire, Etienne
Published: (2025)
Collaborative Group-Aware Hashing for Fast Recommender Systems
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
RAG-Match: Retrieval-Augmented Knowledge Injection and Hierarchical Reasoning for Calibrated Semantic Relevance
by: Jiang, Hengjun, et al.
Published: (2026)
by: Jiang, Hengjun, et al.
Published: (2026)
Prospects of Retrieval Augmented Generation (RAG) for Academic Library Search and Retrieval
by: Bevara, Ravi, et al.
Published: (2025)
by: Bevara, Ravi, et al.
Published: (2025)
DMQR-RAG: Diverse Multi-Query Rewriting for RAG
by: Li, Zhicong, et al.
Published: (2024)
by: Li, Zhicong, et al.
Published: (2024)
CROSSAN: Towards Efficient and Effective Adaptation of Multiple Multimodal Foundation Models for Sequential Recommendation
by: Fu, Junchen, et al.
Published: (2025)
by: Fu, Junchen, et al.
Published: (2025)
FastInsight: Fast and Insightful Retrieval via Fusion Operators for Graph RAG
by: An, Seonho, et al.
Published: (2026)
by: An, Seonho, et al.
Published: (2026)
LARAG: Link-Aware Retrieval Strategy for RAG Systems in Hyperlinked Technical Documentation
by: Bolognesi, Giorgia, et al.
Published: (2026)
by: Bolognesi, Giorgia, et al.
Published: (2026)
A Multi-tiered Solution for Personalized Baggage Item Recommendations using FastText and Association Rule Mining
by: Ravi, Mudavath, et al.
Published: (2025)
by: Ravi, Mudavath, et al.
Published: (2025)
Context-augmented Retrieval: A Novel Framework for Fast Information Retrieval based Response Generation using Large Language Model
by: Ganesh, Sai, et al.
Published: (2024)
by: Ganesh, Sai, et al.
Published: (2024)
LightRAG: Simple and Fast Retrieval-Augmented Generation
by: Guo, Zirui, et al.
Published: (2024)
by: Guo, Zirui, et al.
Published: (2024)
Teach Me How to Denoise: A Universal Framework for Denoising Multi-modal Recommender Systems via Guided Calibration
by: Li, Hongji, et al.
Published: (2025)
by: Li, Hongji, et al.
Published: (2025)
Doctor-RAG: Failure-Aware Repair for Agentic Retrieval-Augmented Generation
by: Jiao, Shuguang, et al.
Published: (2026)
by: Jiao, Shuguang, et al.
Published: (2026)
HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems
by: Tan, Jiejun, et al.
Published: (2024)
by: Tan, Jiejun, et al.
Published: (2024)
Cog-RAG: Cognitive-Inspired Dual-Hypergraph with Theme Alignment Retrieval-Augmented Generation
by: Hu, Hao, et al.
Published: (2025)
by: Hu, Hao, et al.
Published: (2025)
TrackRec: Iterative Alternating Feedback with Chain-of-Thought via Preference Alignment for Recommendation
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
Argo: Efficient Importance Labeling for Enterprise Email Systems
by: Ray, Siddhant, et al.
Published: (2026)
by: Ray, Siddhant, et al.
Published: (2026)
LiteSemRAG: Lightweight LLM-Free Semantic-Aware Graph Retrieval for Robust RAG
by: Yue, Xiao, et al.
Published: (2026)
by: Yue, Xiao, et al.
Published: (2026)
Drift-Aware Continual Tokenization for Generative Recommendation
by: Feng, Yuebo, et al.
Published: (2026)
by: Feng, Yuebo, et al.
Published: (2026)
Do We Still Need GraphRAG? Benchmarking RAG and GraphRAG for Agentic Search Systems
by: Fan, Dongzhe, et al.
Published: (2026)
by: Fan, Dongzhe, et al.
Published: (2026)
Trust or Abstain? A Self-Aware RAG Approach
by: Zhu, Xi, et al.
Published: (2026)
by: Zhu, Xi, et al.
Published: (2026)
OntologyRAG: Better and Faster Biomedical Code Mapping with Retrieval-Augmented Generation (RAG) Leveraging Ontology Knowledge Graphs and Large Language Models
by: Feng, Hui, et al.
Published: (2025)
by: Feng, Hui, et al.
Published: (2025)
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems
by: de Lima, Rafael Teixeira, et al.
Published: (2024)
by: de Lima, Rafael Teixeira, et al.
Published: (2024)
REAPER: Reasoning based Retrieval Planning for Complex RAG Systems
by: Joshi, Ashutosh, et al.
Published: (2024)
by: Joshi, Ashutosh, et al.
Published: (2024)
Towards End-to-End Model-Agnostic Explanations for RAG Systems
by: Sudhi, Viju, et al.
Published: (2025)
by: Sudhi, Viju, et al.
Published: (2025)
EncouRAGe: Evaluating RAG Local, Fast, and Reliable
by: Strich, Jan, et al.
Published: (2025)
by: Strich, Jan, et al.
Published: (2025)
HaS: Accelerating RAG through Homology-Aware Speculative Retrieval
by: Peng, Peng, et al.
Published: (2026)
by: Peng, Peng, et al.
Published: (2026)
R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems
by: Gu, Hao, et al.
Published: (2025)
by: Gu, Hao, et al.
Published: (2025)
Semantic Certainty Assessment in Vector Retrieval Systems: A Novel Framework for Embedding Quality Evaluation
by: Du, Y.
Published: (2025)
by: Du, Y.
Published: (2025)
An Empirical Study of Training ID-Agnostic Multi-modal Sequential Recommenders
by: Li, Youhua, et al.
Published: (2024)
by: Li, Youhua, et al.
Published: (2024)
Similar Items
-
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
by: Feng, Shaoting, et al.
Published: (2025) -
OneAdapt: Fast Configuration Adaptation for Video Analytics Applications via Backpropagation
by: Du, Kuntai, et al.
Published: (2023) -
LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts
by: Gu, Zhuohan, et al.
Published: (2024) -
CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
by: Yao, Jiayi, et al.
Published: (2024) -
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
by: Feng, Shaoting, et al.
Published: (2025)