CaGR-RAG: Context-aware Query Grouping for Disk-based Vector Search in RAG Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeong, Yeonwoo, Park, Kyuli, Cho, Hyunji, Park, Sungyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On The Reproducibility Limitations of RAG Systems
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
von: Song, Jaeyong, et al.
Veröffentlicht: (2026)
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
von: Li, Jiaxing, et al.
Veröffentlicht: (2024)
von: Li, Jiaxing, et al.
Veröffentlicht: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
von: Li, Maoliang, et al.
Veröffentlicht: (2026)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
von: Hong, Guihang, et al.
Veröffentlicht: (2025)
von: Hong, Guihang, et al.
Veröffentlicht: (2025)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
von: Jeong, Bodon, et al.
Veröffentlicht: (2026)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
von: Hwang, Soon, et al.
Veröffentlicht: (2025)
von: Hwang, Soon, et al.
Veröffentlicht: (2025)
Towards Efficient and Scalable Distributed Vector Search with RDMA
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhi, Xiangyu, et al.
Veröffentlicht: (2025)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
von: Kim, Sukjin, et al.
Veröffentlicht: (2025)
von: Kim, Sukjin, et al.
Veröffentlicht: (2025)
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
von: Addison, Parker, et al.
Veröffentlicht: (2024)
von: Addison, Parker, et al.
Veröffentlicht: (2024)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
von: Gui, Yuntao, et al.
Veröffentlicht: (2025)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
von: Xu, Yuming, et al.
Veröffentlicht: (2025)
von: Xu, Yuming, et al.
Veröffentlicht: (2025)
SQUASH: Serverless and Distributed Quantization-based Attributed Vector Similarity Search
von: Oakley, Joe, et al.
Veröffentlicht: (2025)
von: Oakley, Joe, et al.
Veröffentlicht: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
von: Park, Gijun
Veröffentlicht: (2026)
von: Park, Gijun
Veröffentlicht: (2026)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
von: Liu, Yi, et al.
Veröffentlicht: (2025)
von: Liu, Yi, et al.
Veröffentlicht: (2025)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
von: Maharaj, Davendra, et al.
Veröffentlicht: (2026)
von: Maharaj, Davendra, et al.
Veröffentlicht: (2026)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
von: Lin, Chien-Yu, et al.
Veröffentlicht: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
von: Hwang, Jinwoo, et al.
Veröffentlicht: (2025)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2025)
von: Rashid, Sulaiman Muhammad, et al.
Veröffentlicht: (2025)
A Hierarchical Sharded Blockchain Balancing Performance and Availability
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
von: Jo, Yongrae, et al.
Veröffentlicht: (2025)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
von: Chai, Huichao, et al.
Veröffentlicht: (2026)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2025)
von: Cho, Jaehong, et al.
Veröffentlicht: (2025)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
von: Ranjan, Desh, et al.
Veröffentlicht: (2024)
von: Ranjan, Desh, et al.
Veröffentlicht: (2024)
vPALs: Towards Verified Performance-aware Learning System For Resource Management
von: He, Guoliang, et al.
Veröffentlicht: (2024)
von: He, Guoliang, et al.
Veröffentlicht: (2024)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
von: Wei, Wei, et al.
Veröffentlicht: (2024)
von: Wei, Wei, et al.
Veröffentlicht: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
von: Adikari, Tharindu, et al.
Veröffentlicht: (2024)
von: Adikari, Tharindu, et al.
Veröffentlicht: (2024)
Parallel R-tree-based Spatial Query Processing on a Commercial Processing-in-Memory System
von: Jannat, Tasmia, et al.
Veröffentlicht: (2026)
von: Jannat, Tasmia, et al.
Veröffentlicht: (2026)
Multi-Objective Optimization of Consumer Group Autoscaling in Message Broker Systems
von: Landau, Diogo, et al.
Veröffentlicht: (2024)
von: Landau, Diogo, et al.
Veröffentlicht: (2024)
A Two-Level Thermal Cycling-aware Task Mapping Technique for Reliability Management in Manycore Systems
von: Khani, Fatemeh Hossein, et al.
Veröffentlicht: (2024)
von: Khani, Fatemeh Hossein, et al.
Veröffentlicht: (2024)
Search for shortest paths based on a projective description of unweighted graphs
von: Melent'ev, V. A.
Veröffentlicht: (2024)
von: Melent'ev, V. A.
Veröffentlicht: (2024)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
von: Cho, Jaehong, et al.
Veröffentlicht: (2026)
Affinity-aware Serverless Function Scheduling
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
von: De Palma, Giuseppe, et al.
Veröffentlicht: (2024)
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving System
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
von: Bai, Fengyao, et al.
Veröffentlicht: (2026)
Vectorized Sequence-Based Chunking for Data Deduplication
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
von: Udayashankar, Sreeharsha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On The Reproducibility Limitations of RAG Systems
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025) -
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
von: Song, Jaeyong, et al.
Veröffentlicht: (2026) -
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
von: Li, Jiaxing, et al.
Veröffentlicht: (2024) -
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
von: Wang, Wenfeng, et al.
Veröffentlicht: (2026) -
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
von: Liu, Kaiwei, et al.
Veröffentlicht: (2025)