MobileRAG: A Fast, Memory-Efficient, and Energy-Efficient Method for On-Device RAG
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Park, Taehwan, Lee, Geonho, Kim, Min-Soo |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ExtGraph: A Fast Extraction Method of User-intended Graphs from a Relational Database
par: Park, Jeongho, et autres
Publié: (2025)
par: Park, Jeongho, et autres
Publié: (2025)
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
par: Loo, Gowen, et autres
Publié: (2025)
par: Loo, Gowen, et autres
Publié: (2025)
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
par: Yan, Jianxin, et autres
Publié: (2026)
par: Yan, Jianxin, et autres
Publié: (2026)
Cost-Efficient RAG for Entity Matching with LLMs: A Blocking-based Exploration
par: Ma, Chuangtao, et autres
Publié: (2026)
par: Ma, Chuangtao, et autres
Publié: (2026)
cuRPQ: A High-Performance GPU-Based Framework for Processing Regular and Conjunctive Regular Path Queries
par: Park, Sungwoo, et autres
Publié: (2026)
par: Park, Sungwoo, et autres
Publié: (2026)
EvoRAG: Making Knowledge Graph-based RAG Automatically Evolve through Feedback-driven Backpropagation
par: Fu, Zhenbo, et autres
Publié: (2026)
par: Fu, Zhenbo, et autres
Publié: (2026)
HedraRAG: Coordinating LLM Generation and Database Retrieval in Heterogeneous RAG Serving
par: Hu, Zhengding, et autres
Publié: (2025)
par: Hu, Zhengding, et autres
Publié: (2025)
RAGPulse: An Open-Source RAG Workload Trace to Optimize RAG Serving Systems
par: Wang, Zhengchao, et autres
Publié: (2025)
par: Wang, Zhengchao, et autres
Publié: (2025)
RAG-Stack: Co-Optimizing RAG Quality and Performance From the Vector Database Perspective
par: Jiang, Wenqi
Publié: (2025)
par: Jiang, Wenqi
Publié: (2025)
SEFRQO: A Self-Evolving Fine-Tuned RAG-Based Query Optimizer
par: Liu, Hanwen, et autres
Publié: (2025)
par: Liu, Hanwen, et autres
Publié: (2025)
GateANN: I/O-Efficient Filtered Vector Search on SSDs
par: Lee, Nakyung, et autres
Publié: (2026)
par: Lee, Nakyung, et autres
Publié: (2026)
Multi-Meta-RAG: Improving RAG for Multi-Hop Queries using Database Filtering with LLM-Extracted Metadata
par: Poliakov, Mykhailo, et autres
Publié: (2024)
par: Poliakov, Mykhailo, et autres
Publié: (2024)
Energy-Efficient Path Planning with Multi-Location Object Pickup for Mobile Robots on Uneven Terrain
par: Babakano, Faiza, et autres
Publié: (2025)
par: Babakano, Faiza, et autres
Publié: (2025)
RAG-Driven Data Quality Governance for Enterprise ERP Systems
par: Vedat, Sedat Bin, et autres
Publié: (2025)
par: Vedat, Sedat Bin, et autres
Publié: (2025)
RELOAD: A Robust and Efficient Learned Query Optimizer for Database Systems
par: Lee, Seokwon, et autres
Publié: (2026)
par: Lee, Seokwon, et autres
Publié: (2026)
SAGE: A Framework of Precise Retrieval for RAG
par: Zhang, Jintao, et autres
Publié: (2025)
par: Zhang, Jintao, et autres
Publié: (2025)
Beyond Quacking: Deep Integration of Language Models and RAG into DuckDB
par: Dorbani, Anas, et autres
Publié: (2025)
par: Dorbani, Anas, et autres
Publié: (2025)
Efficient Vector Search on Disaggregated Memory with d-HNSW
par: Liu, Yi, et autres
Publié: (2025)
par: Liu, Yi, et autres
Publié: (2025)
PentaRAG: Large-Scale Intelligent Knowledge Retrieval for Enterprise LLM Applications
par: Syarubany, Abu Hanif Muhammad, et autres
Publié: (2025)
par: Syarubany, Abu Hanif Muhammad, et autres
Publié: (2025)
CARROT: A Learned Cost-Constrained Retrieval Optimization System for RAG
par: Wang, Ziting, et autres
Publié: (2024)
par: Wang, Ziting, et autres
Publié: (2024)
Balancing Content Size in RAG-Text2SQL System
par: Gurawa, Prakhar, et autres
Publié: (2025)
par: Gurawa, Prakhar, et autres
Publié: (2025)
In-depth Analysis of Graph-based RAG in a Unified Framework
par: Zhou, Yingli, et autres
Publié: (2025)
par: Zhou, Yingli, et autres
Publié: (2025)
Needle-in-RAG: Prompt-Conditioned Character-Level Traceback of Poisoned Spans in Retrieved Evidence
par: Cui, Huining, et autres
Publié: (2026)
par: Cui, Huining, et autres
Publié: (2026)
BubbleRAG: Evidence-Driven Retrieval-Augmented Generation for Black-Box Knowledge Graphs
par: Pan, Duyi, et autres
Publié: (2026)
par: Pan, Duyi, et autres
Publié: (2026)
FastInsight: Fast and Insightful Retrieval via Fusion Operators for Graph RAG
par: An, Seonho, et autres
Publié: (2026)
par: An, Seonho, et autres
Publié: (2026)
Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective
par: Yang, Rui, et autres
Publié: (2024)
par: Yang, Rui, et autres
Publié: (2024)
PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
par: Luo, Jigao, et autres
Publié: (2025)
par: Luo, Jigao, et autres
Publié: (2025)
LatentTune: Efficient Tuning of High Dimensional Database Parameters via Latent Representation Learning
par: Kwon, Sein, et autres
Publié: (2026)
par: Kwon, Sein, et autres
Publié: (2026)
A2RAG: Adaptive Agentic Graph Retrieval for Cost-Aware and Reliable Reasoning
par: Liu, Jiate, et autres
Publié: (2026)
par: Liu, Jiate, et autres
Publié: (2026)
PairwiseHist: Fast, Accurate and Space-Efficient Approximate Query Processing with Data Compression
par: Hurst, Aaron, et autres
Publié: (2024)
par: Hurst, Aaron, et autres
Publié: (2024)
Optimization of embeddings storage for RAG systems using quantization and dimensionality reduction techniques
par: Huerga-Pérez, Naamán, et autres
Publié: (2025)
par: Huerga-Pérez, Naamán, et autres
Publié: (2025)
Chipmink: Efficient Delta Identification for Massive Object Graph
par: Chockchowwat, Supawit, et autres
Publié: (2025)
par: Chockchowwat, Supawit, et autres
Publié: (2025)
Efficient Query Repair for Aggregate Constraints
par: Algarni, Shatha, et autres
Publié: (2025)
par: Algarni, Shatha, et autres
Publié: (2025)
DLHT: A Non-blocking Resizable Hashtable with Fast Deletes and Memory-awareness
par: Katsarakis, Antonios, et autres
Publié: (2024)
par: Katsarakis, Antonios, et autres
Publié: (2024)
LHGstore: An In-Memory Learned Graph Storage for Fast Updates and Analytics
par: Qiao, Pengpeng, et autres
Publié: (2026)
par: Qiao, Pengpeng, et autres
Publié: (2026)
Efficient Defective Clique Enumeration and Search with Worst-Case Optimal Search Space
par: Jang, Jihoon, et autres
Publié: (2025)
par: Jang, Jihoon, et autres
Publié: (2025)
QCore: Data-Efficient, On-Device Continual Calibration for Quantized Models -- Extended Version
par: Campos, David, et autres
Publié: (2024)
par: Campos, David, et autres
Publié: (2024)
DIST: Efficient k-Clique Listing via Induced Subgraph Trie
par: Nam, Yehyun, et autres
Publié: (2025)
par: Nam, Yehyun, et autres
Publié: (2025)
A Memory-Efficient Distributed Algorithm for Approximate Nearest Neighbour Search with Arbitrary Distances
par: Garcia-Morato, Elena, et autres
Publié: (2024)
par: Garcia-Morato, Elena, et autres
Publié: (2024)
LEGO-GraphRAG: Modularizing Graph-based Retrieval-Augmented Generation for Design Space Exploration
par: Cao, Yukun, et autres
Publié: (2024)
par: Cao, Yukun, et autres
Publié: (2024)
Documents similaires
-
ExtGraph: A Fast Extraction Method of User-intended Graphs from a Relational Database
par: Park, Jeongho, et autres
Publié: (2025) -
MobileRAG: Enhancing Mobile Agent with Retrieval-Augmented Generation
par: Loo, Gowen, et autres
Publié: (2025) -
QCFuse: Query-Centric Cache Fusion for Efficient RAG Inference
par: Yan, Jianxin, et autres
Publié: (2026) -
Cost-Efficient RAG for Entity Matching with LLMs: A Blocking-based Exploration
par: Ma, Chuangtao, et autres
Publié: (2026) -
cuRPQ: A High-Performance GPU-Based Framework for Processing Regular and Conjunctive Regular Path Queries
par: Park, Sungwoo, et autres
Publié: (2026)