CaGR-RAG: Context-aware Query Grouping for Disk-based Vector Search in RAG Systems
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Jeong, Yeonwoo, Park, Kyuli, Cho, Hyunji, Park, Sungyong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
On The Reproducibility Limitations of RAG Systems
par: Wang, Baiqiang, et autres
Publié: (2025)
par: Wang, Baiqiang, et autres
Publié: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
par: Song, Jaeyong, et autres
Publié: (2026)
par: Song, Jaeyong, et autres
Publié: (2026)
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
par: Li, Jiaxing, et autres
Publié: (2024)
par: Li, Jiaxing, et autres
Publié: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
par: Wang, Wenfeng, et autres
Publié: (2026)
par: Wang, Wenfeng, et autres
Publié: (2026)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
par: Liu, Kaiwei, et autres
Publié: (2025)
par: Liu, Kaiwei, et autres
Publié: (2025)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
par: Yu, Weiping, et autres
Publié: (2025)
par: Yu, Weiping, et autres
Publié: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
par: Li, Maoliang, et autres
Publié: (2026)
par: Li, Maoliang, et autres
Publié: (2026)
Accelerating LLM Inference with Precomputed Query Storage
par: Park, Jay H., et autres
Publié: (2025)
par: Park, Jay H., et autres
Publié: (2025)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
par: Hong, Guihang, et autres
Publié: (2025)
par: Hong, Guihang, et autres
Publié: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
par: Zhang, Huawei, et autres
Publié: (2025)
par: Zhang, Huawei, et autres
Publié: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
par: Jeong, Bodon, et autres
Publié: (2026)
par: Jeong, Bodon, et autres
Publié: (2026)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
par: Hwang, Soon, et autres
Publié: (2025)
par: Hwang, Soon, et autres
Publié: (2025)
Towards Efficient and Scalable Distributed Vector Search with RDMA
par: Zhi, Xiangyu, et autres
Publié: (2025)
par: Zhi, Xiangyu, et autres
Publié: (2025)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
par: Kim, Sukjin, et autres
Publié: (2025)
par: Kim, Sukjin, et autres
Publié: (2025)
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
par: Addison, Parker, et autres
Publié: (2024)
par: Addison, Parker, et autres
Publié: (2024)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
par: Gui, Yuntao, et autres
Publié: (2025)
par: Gui, Yuntao, et autres
Publié: (2025)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
par: Xu, Yuming, et autres
Publié: (2025)
par: Xu, Yuming, et autres
Publié: (2025)
SQUASH: Serverless and Distributed Quantization-based Attributed Vector Similarity Search
par: Oakley, Joe, et autres
Publié: (2025)
par: Oakley, Joe, et autres
Publié: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
par: Park, Gijun
Publié: (2026)
par: Park, Gijun
Publié: (2026)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
par: Liu, Yi, et autres
Publié: (2025)
par: Liu, Yi, et autres
Publié: (2025)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
par: Miyashita, Yusuke, et autres
Publié: (2024)
par: Miyashita, Yusuke, et autres
Publié: (2024)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
par: Maharaj, Davendra, et autres
Publié: (2026)
par: Maharaj, Davendra, et autres
Publié: (2026)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
par: Lin, Chien-Yu, et autres
Publié: (2025)
par: Lin, Chien-Yu, et autres
Publié: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
par: Hwang, Jinwoo, et autres
Publié: (2025)
par: Hwang, Jinwoo, et autres
Publié: (2025)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
par: Rashid, Sulaiman Muhammad, et autres
Publié: (2025)
par: Rashid, Sulaiman Muhammad, et autres
Publié: (2025)
A Hierarchical Sharded Blockchain Balancing Performance and Availability
par: Jo, Yongrae, et autres
Publié: (2025)
par: Jo, Yongrae, et autres
Publié: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
par: Chai, Huichao, et autres
Publié: (2026)
par: Chai, Huichao, et autres
Publié: (2026)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
par: Cho, Jaehong, et autres
Publié: (2025)
par: Cho, Jaehong, et autres
Publié: (2025)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
par: Ranjan, Desh, et autres
Publié: (2024)
par: Ranjan, Desh, et autres
Publié: (2024)
vPALs: Towards Verified Performance-aware Learning System For Resource Management
par: He, Guoliang, et autres
Publié: (2024)
par: He, Guoliang, et autres
Publié: (2024)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
par: Wei, Wei, et autres
Publié: (2024)
par: Wei, Wei, et autres
Publié: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
par: Adikari, Tharindu, et autres
Publié: (2024)
par: Adikari, Tharindu, et autres
Publié: (2024)
Parallel R-tree-based Spatial Query Processing on a Commercial Processing-in-Memory System
par: Jannat, Tasmia, et autres
Publié: (2026)
par: Jannat, Tasmia, et autres
Publié: (2026)
Multi-Objective Optimization of Consumer Group Autoscaling in Message Broker Systems
par: Landau, Diogo, et autres
Publié: (2024)
par: Landau, Diogo, et autres
Publié: (2024)
A Two-Level Thermal Cycling-aware Task Mapping Technique for Reliability Management in Manycore Systems
par: Khani, Fatemeh Hossein, et autres
Publié: (2024)
par: Khani, Fatemeh Hossein, et autres
Publié: (2024)
Search for shortest paths based on a projective description of unweighted graphs
par: Melent'ev, V. A.
Publié: (2024)
par: Melent'ev, V. A.
Publié: (2024)
Affinity-aware Serverless Function Scheduling
par: De Palma, Giuseppe, et autres
Publié: (2024)
par: De Palma, Giuseppe, et autres
Publié: (2024)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
par: Cho, Jaehong, et autres
Publié: (2026)
par: Cho, Jaehong, et autres
Publié: (2026)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
par: Bai, Fengyao, et autres
Publié: (2026)
par: Bai, Fengyao, et autres
Publié: (2026)
Vectorized Sequence-Based Chunking for Data Deduplication
par: Udayashankar, Sreeharsha, et autres
Publié: (2025)
par: Udayashankar, Sreeharsha, et autres
Publié: (2025)
Documents similaires
-
On The Reproducibility Limitations of RAG Systems
par: Wang, Baiqiang, et autres
Publié: (2025) -
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
par: Song, Jaeyong, et autres
Publié: (2026) -
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
par: Li, Jiaxing, et autres
Publié: (2024) -
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
par: Wang, Wenfeng, et autres
Publié: (2026) -
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
par: Liu, Kaiwei, et autres
Publié: (2025)