Saved in:
| Main Authors: | Jeong, Yeonwoo, Park, Kyuli, Cho, Hyunji, Park, Sungyong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.01164 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On The Reproducibility Limitations of RAG Systems
by: Wang, Baiqiang, et al.
Published: (2025)
by: Wang, Baiqiang, et al.
Published: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
by: Li, Jiaxing, et al.
Published: (2024)
by: Li, Jiaxing, et al.
Published: (2024)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)
by: Jeong, Bodon, et al.
Published: (2026)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
by: Yu, Weiping, et al.
Published: (2025)
by: Yu, Weiping, et al.
Published: (2025)
Accelerating LLM Inference with Precomputed Query Storage
by: Park, Jay H., et al.
Published: (2025)
by: Park, Jay H., et al.
Published: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
by: Liu, Kaiwei, et al.
Published: (2025)
by: Liu, Kaiwei, et al.
Published: (2025)
HeRo: Adaptive Orchestration of Agentic RAG on Heterogeneous Mobile SoC
by: Li, Maoliang, et al.
Published: (2026)
by: Li, Maoliang, et al.
Published: (2026)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
by: Zhang, Huawei, et al.
Published: (2025)
by: Zhang, Huawei, et al.
Published: (2025)
OASIS: Object-based Analytics Storage for Intelligent SQL Query Offloading in Scientific Tabular Workloads
by: Hwang, Soon, et al.
Published: (2025)
by: Hwang, Soon, et al.
Published: (2025)
CoEdge-RAG: Optimizing Hierarchical Scheduling for Retrieval-Augmented LLMs in Collaborative Edge Computing
by: Hong, Guihang, et al.
Published: (2025)
by: Hong, Guihang, et al.
Published: (2025)
C-FedRAG: A Confidential Federated Retrieval-Augmented Generation System
by: Addison, Parker, et al.
Published: (2024)
by: Addison, Parker, et al.
Published: (2024)
Towards Efficient and Scalable Distributed Vector Search with RDMA
by: Zhi, Xiangyu, et al.
Published: (2025)
by: Zhi, Xiangyu, et al.
Published: (2025)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
by: Kim, Sukjin, et al.
Published: (2025)
by: Kim, Sukjin, et al.
Published: (2025)
Déjà Vu: Efficient Video-Language Query Engine with Learning-based Inter-Frame Computation Reuse
by: Hwang, Jinwoo, et al.
Published: (2025)
by: Hwang, Jinwoo, et al.
Published: (2025)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
by: Gui, Yuntao, et al.
Published: (2025)
by: Gui, Yuntao, et al.
Published: (2025)
SQUASH: Serverless and Distributed Quantization-based Attributed Vector Similarity Search
by: Oakley, Joe, et al.
Published: (2025)
by: Oakley, Joe, et al.
Published: (2025)
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
by: Xu, Yuming, et al.
Published: (2025)
by: Xu, Yuming, et al.
Published: (2025)
Quantifying Autoscaler Vulnerabilities: An Empirical Study of Resource Misallocation Induced by Cloud Infrastructure Faults
by: Park, Gijun
Published: (2026)
by: Park, Gijun
Published: (2026)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
by: Miyashita, Yusuke, et al.
Published: (2024)
by: Miyashita, Yusuke, et al.
Published: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure
by: Cho, Jaehong, et al.
Published: (2025)
by: Cho, Jaehong, et al.
Published: (2025)
DGNNFlow: A Streaming Dataflow Architecture for Real-Time Edge-based Dynamic GNN Inference in HL-LHC Trigger Systems
by: Maharaj, Davendra, et al.
Published: (2026)
by: Maharaj, Davendra, et al.
Published: (2026)
TurboGR: An Accelerated Training System for Large-Scale Generative Recommendation
by: Chai, Huichao, et al.
Published: (2026)
by: Chai, Huichao, et al.
Published: (2026)
Distributed Hierarchical Machine Learning for Joint Resource Allocation and Slice Selection in In-Network Edge Systems
by: Rashid, Sulaiman Muhammad, et al.
Published: (2025)
by: Rashid, Sulaiman Muhammad, et al.
Published: (2025)
A Hierarchical Sharded Blockchain Balancing Performance and Availability
by: Jo, Yongrae, et al.
Published: (2025)
by: Jo, Yongrae, et al.
Published: (2025)
LLMServingSim 2.0: A Unified Simulator for Heterogeneous and Disaggregated LLM Serving Infrastructure
by: Cho, Jaehong, et al.
Published: (2026)
by: Cho, Jaehong, et al.
Published: (2026)
Parallel R-tree-based Spatial Query Processing on a Commercial Processing-in-Memory System
by: Jannat, Tasmia, et al.
Published: (2026)
by: Jannat, Tasmia, et al.
Published: (2026)
An AD based library for Efficient Hessian and Hessian-Vector Product Computation on GPU
by: Ranjan, Desh, et al.
Published: (2024)
by: Ranjan, Desh, et al.
Published: (2024)
vPALs: Towards Verified Performance-aware Learning System For Resource Management
by: He, Guoliang, et al.
Published: (2024)
by: He, Guoliang, et al.
Published: (2024)
Energy-aware Incremental OTA Update for Flash-based Batteryless IoT Devices
by: Wei, Wei, et al.
Published: (2024)
by: Wei, Wei, et al.
Published: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
by: Adikari, Tharindu, et al.
Published: (2024)
by: Adikari, Tharindu, et al.
Published: (2024)
LLMServingSim: A HW/SW Co-Simulation Infrastructure for LLM Inference Serving at Scale
by: Cho, Jaehong, et al.
Published: (2024)
by: Cho, Jaehong, et al.
Published: (2024)
Multi-Objective Optimization of Consumer Group Autoscaling in Message Broker Systems
by: Landau, Diogo, et al.
Published: (2024)
by: Landau, Diogo, et al.
Published: (2024)
A Two-Level Thermal Cycling-aware Task Mapping Technique for Reliability Management in Manycore Systems
by: Khani, Fatemeh Hossein, et al.
Published: (2024)
by: Khani, Fatemeh Hossein, et al.
Published: (2024)
DFLOP: A Data-driven Framework for Multimodal LLM Training Pipeline Optimization
by: An, Hyeonjun, et al.
Published: (2026)
by: An, Hyeonjun, et al.
Published: (2026)
Breaking the Capacity Bottleneck in Model-Heterogeneous Federated Learning via Gradual Model Restoration
by: Ma, Chengjie, et al.
Published: (2025)
by: Ma, Chengjie, et al.
Published: (2025)
NestedFP: High-Performance, Memory-Efficient Dual-Precision Floating Point Support for LLMs
by: Lee, Haeun, et al.
Published: (2025)
by: Lee, Haeun, et al.
Published: (2025)
Similar Items
-
On The Reproducibility Limitations of RAG Systems
by: Wang, Baiqiang, et al.
Published: (2025) -
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026) -
EACO-RAG: Towards Distributed Tiered LLM Deployment using Edge-Assisted and Collaborative RAG with Adaptive Knowledge Update
by: Li, Jiaxing, et al.
Published: (2024) -
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026) -
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference
by: Jeong, Bodon, et al.
Published: (2026)