Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Search
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Dong, Yu, Yanxuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
by: Liu, Dong, et al.
Published: (2026)
by: Liu, Dong, et al.
Published: (2026)
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
by: Tan, Zhiwen, et al.
Published: (2025)
by: Tan, Zhiwen, et al.
Published: (2025)
Breaking the Autoregressive Chain: Hyper-Parallel Decoding for Efficient LLM-Based Attribute Value Extraction
by: Glavas, Theodore, et al.
Published: (2026)
by: Glavas, Theodore, et al.
Published: (2026)
TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
by: Liu, Dong, et al.
Published: (2025)
by: Liu, Dong, et al.
Published: (2025)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
by: Zhuang, Ziyuan, et al.
Published: (2024)
by: Zhuang, Ziyuan, et al.
Published: (2024)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
by: Zhang, Mozhi, et al.
Published: (2025)
by: Zhang, Mozhi, et al.
Published: (2025)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
by: Fu, Jia, et al.
Published: (2024)
by: Fu, Jia, et al.
Published: (2024)
Parallel Key-Value Cache Fusion for Position Invariant RAG
by: Oh, Philhoon, et al.
Published: (2025)
by: Oh, Philhoon, et al.
Published: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
by: Dong, Sicheng, et al.
Published: (2025)
by: Dong, Sicheng, et al.
Published: (2025)
Towards Better Monolingual Japanese Retrievers with Multi-Vector Models
by: Clavié, Benjamin
Published: (2023)
by: Clavié, Benjamin
Published: (2023)
Entropic Claim Resolution: Uncertainty-Driven Evidence Selection for RAG
by: Di Gioia, Davide
Published: (2026)
by: Di Gioia, Davide
Published: (2026)
ParallelSearch: Train your LLMs to Decompose Query and Search Sub-queries in Parallel with Reinforcement Learning
by: Zhao, Shu, et al.
Published: (2025)
by: Zhao, Shu, et al.
Published: (2025)
PA-RAG: RAG Alignment via Multi-Perspective Preference Optimization
by: Wu, Jiayi, et al.
Published: (2024)
by: Wu, Jiayi, et al.
Published: (2024)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
Dynamic Task Vector Grouping for Efficient Multi-Task Prompt Tuning
by: Zhang, Pieyi, et al.
Published: (2025)
by: Zhang, Pieyi, et al.
Published: (2025)
SFR-RAG: Towards Contextually Faithful LLMs
by: Nguyen, Xuan-Phi, et al.
Published: (2024)
by: Nguyen, Xuan-Phi, et al.
Published: (2024)
HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
by: Liu, Pei, et al.
Published: (2025)
by: Liu, Pei, et al.
Published: (2025)
Skeleton-of-Thought: Prompting LLMs for Efficient Parallel Generation
by: Ning, Xuefei, et al.
Published: (2023)
by: Ning, Xuefei, et al.
Published: (2023)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
by: BehnamGhader, Parishad, et al.
Published: (2024)
by: BehnamGhader, Parishad, et al.
Published: (2024)
PropRAG: Guiding Retrieval with Beam Search over Proposition Paths
by: Wang, Jingjin, et al.
Published: (2025)
by: Wang, Jingjin, et al.
Published: (2025)
Beyond Component Strength: Synergistic Integration and Adaptive Calibration in Multi-Agent RAG Systems
by: Krishnan, Jithin
Published: (2025)
by: Krishnan, Jithin
Published: (2025)
Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning
by: Lin, Junhong, et al.
Published: (2025)
by: Lin, Junhong, et al.
Published: (2025)
HalluSearch at SemEval-2025 Task 3: A Search-Enhanced RAG Pipeline for Hallucination Detection
by: Abdallah, Mohamed A., et al.
Published: (2025)
by: Abdallah, Mohamed A., et al.
Published: (2025)
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025)
by: Chawla, Kushal, et al.
Published: (2025)
Balancing Rewards in Text Summarization: Multi-Objective Reinforcement Learning via HyperVolume Optimization
by: Song, Junjie, et al.
Published: (2025)
by: Song, Junjie, et al.
Published: (2025)
SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG
by: Si, Xiaonan, et al.
Published: (2025)
by: Si, Xiaonan, et al.
Published: (2025)
MES-RAG: Bringing Multi-modal, Entity-Storage, and Secure Enhancements to RAG
by: Wu, Pingyu, et al.
Published: (2025)
by: Wu, Pingyu, et al.
Published: (2025)
RETVec: Resilient and Efficient Text Vectorizer
by: Bursztein, Elie, et al.
Published: (2023)
by: Bursztein, Elie, et al.
Published: (2023)
ComRAG: Retrieval-Augmented Generation with Dynamic Vector Stores for Real-time Community Question Answering in Industry
by: Chen, Qinwen, et al.
Published: (2025)
by: Chen, Qinwen, et al.
Published: (2025)
Cerberus: Efficient Inference with Adaptive Parallel Decoding and Sequential Knowledge Enhancement
by: Liu, Yuxuan, et al.
Published: (2024)
by: Liu, Yuxuan, et al.
Published: (2024)
MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
by: Nguyen, Thang, et al.
Published: (2025)
by: Nguyen, Thang, et al.
Published: (2025)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
by: Zibakhsh, Soheil, et al.
Published: (2025)
by: Zibakhsh, Soheil, et al.
Published: (2025)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
UniMS-RAG: A Unified Multi-source Retrieval-Augmented Generation for Personalized Dialogue Systems
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
Know3-RAG: A Knowledge-aware RAG Framework with Adaptive Retrieval, Generation, and Filtering
by: Liu, Xukai, et al.
Published: (2025)
by: Liu, Xukai, et al.
Published: (2025)
CorIL: Towards Enriching Indian Language to Indian Language Parallel Corpora and Machine Translation Systems
by: Bhattacharjee, Soham, et al.
Published: (2025)
by: Bhattacharjee, Soham, et al.
Published: (2025)
Similar Items
-
$π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modeling
by: Liu, Dong, et al.
Published: (2025) -
SemToken: Semantic-Aware Tokenization for Efficient Long-Context Language Modeling
by: Liu, Dong, et al.
Published: (2025) -
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
by: Liu, Dong, et al.
Published: (2025) -
Thoughts-as-Planning: Latent World Models for Chain-of-Thoughts Optimization via Reinforcement Planning
by: Liu, Dong, et al.
Published: (2026) -
RAG-R1: Incentivizing the Search and Reasoning Capabilities of LLMs through Multi-query Parallelism
by: Tan, Zhiwen, et al.
Published: (2025)