Towards Efficient and Scalable Distributed Vector Search with RDMA
Fuente:
arXiv
Saved in:
| Main Authors: | Zhi, Xiangyu, Chen, Meng, Yan, Xiao, Lu, Baotong, Li, Hui, Zhang, Qianxi, Chen, Qi, Cheng, James |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
by: Xu, Yuming, et al.
Published: (2025)
by: Xu, Yuming, et al.
Published: (2025)
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
by: Gui, Yuntao, et al.
Published: (2025)
by: Gui, Yuntao, et al.
Published: (2025)
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
by: Licker, Nandor, et al.
Published: (2025)
by: Licker, Nandor, et al.
Published: (2025)
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
by: Chen, June, et al.
Published: (2026)
by: Chen, June, et al.
Published: (2026)
ALock: Asymmetric Lock Primitive for RDMA Systems
by: Baran, Amanda, et al.
Published: (2024)
by: Baran, Amanda, et al.
Published: (2024)
Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
by: Brock, Benjamin, et al.
Published: (2023)
by: Brock, Benjamin, et al.
Published: (2023)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
by: Lu, Baotong, et al.
Published: (2024)
by: Lu, Baotong, et al.
Published: (2024)
The Semantic Arrow of Time, Part III: RDMA and the Completion Fallacy
by: Borrill, Paul
Published: (2026)
by: Borrill, Paul
Published: (2026)
Expert-as-a-Service: Towards Efficient, Scalable, and Robust Large-scale MoE Serving
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Towards a Scalable and Efficient PGAS-based Distributed OpenMP
by: Shan, Baodi, et al.
Published: (2024)
by: Shan, Baodi, et al.
Published: (2024)
FedRDMA: Communication-Efficient Cross-Silo Federated LLM via Chunked RDMA Transmission
by: Zhang, Zeling, et al.
Published: (2024)
by: Zhang, Zeling, et al.
Published: (2024)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
Towards the Distributed Large-scale k-NN Graph Construction by Graph Merge
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
Efficient Distributed MLLM Training with Cornstarch
by: Jang, Insu, et al.
Published: (2025)
by: Jang, Insu, et al.
Published: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
Janus: Disaggregating Attention and Experts for Scalable MoE Inference
by: Zhang, Zhexiang, et al.
Published: (2025)
by: Zhang, Zhexiang, et al.
Published: (2025)
FuxiShuffle: An Adaptive and Resilient Shuffle Service for Distributed Data Processing on Alibaba Cloud
by: Lin, Yuhao, et al.
Published: (2026)
by: Lin, Yuhao, et al.
Published: (2026)
SQUASH: Serverless and Distributed Quantization-based Attributed Vector Similarity Search
by: Oakley, Joe, et al.
Published: (2025)
by: Oakley, Joe, et al.
Published: (2025)
Reimagining RDMA Through the Lens of ML
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
Pilotfish: Distributed Execution for Scalable Blockchains
by: Kniep, Quentin, et al.
Published: (2024)
by: Kniep, Quentin, et al.
Published: (2024)
Efficient Multi-Worker Selection based Distributed Swarm Learning via Analog Aggregation
by: Yao, Zhuoyu, et al.
Published: (2025)
by: Yao, Zhuoyu, et al.
Published: (2025)
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
by: Xu, Minxian, et al.
Published: (2026)
by: Xu, Minxian, et al.
Published: (2026)
High-Dimensional Distributed Sparse Classification with Scalable Communication-Efficient Global Updates
by: Lu, Fred, et al.
Published: (2024)
by: Lu, Fred, et al.
Published: (2024)
Handling of Memory Page Faults during Virtual-Address RDMA
by: Psistakis, Antonis
Published: (2025)
by: Psistakis, Antonis
Published: (2025)
Toward Heterogeneous, Distributed, and Energy-Efficient Computing with SYCL
by: Cosenza, Biagio, et al.
Published: (2025)
by: Cosenza, Biagio, et al.
Published: (2025)
OptiNIC: A Resilient and Tail-Optimal RDMA NIC for Distributed ML Workloads
by: Warraich, Ertza, et al.
Published: (2025)
by: Warraich, Ertza, et al.
Published: (2025)
SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods
by: Yang, Shaofeng, et al.
Published: (2026)
by: Yang, Shaofeng, et al.
Published: (2026)
DIMS: Distributed Index for Similarity Search in Metric Spaces
by: Zhu, Yifan, et al.
Published: (2024)
by: Zhu, Yifan, et al.
Published: (2024)
Efficient Distributed Learning over Decentralized Networks with Convoluted Support Vector Machine
by: Chen, Canyi, et al.
Published: (2025)
by: Chen, Canyi, et al.
Published: (2025)
Accelerating Biclique Counting on GPU
by: Qiu, Linshan, et al.
Published: (2024)
by: Qiu, Linshan, et al.
Published: (2024)
gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Varuna: Enabling Failure-Type Aware RDMA Failover
by: Wang, Xiaoyang, et al.
Published: (2026)
by: Wang, Xiaoyang, et al.
Published: (2026)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource Multiplexing
by: Wang, Yanbo, et al.
Published: (2026)
by: Wang, Yanbo, et al.
Published: (2026)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
Towards Resource-Efficient Serverless LLM Inference with SLINFER
by: Xu, Chuhao, et al.
Published: (2025)
by: Xu, Chuhao, et al.
Published: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
by: Song, Jaeyong, et al.
Published: (2026)
by: Song, Jaeyong, et al.
Published: (2026)
CondenseGraph: Communication-Efficient Distributed GNN Training via On-the-Fly Graph Condensation
by: Zhang, Zizhao, et al.
Published: (2026)
by: Zhang, Zizhao, et al.
Published: (2026)
Similar Items
-
Scalable Distributed Vector Search via Accuracy Preserving Index Construction
by: Xu, Yuming, et al.
Published: (2025) -
PilotANN: Memory-Bounded GPU Acceleration for Vector Search
by: Gui, Yuntao, et al.
Published: (2025) -
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
by: Licker, Nandor, et al.
Published: (2025) -
OnePiece: A Large-Scale Distributed Inference System with RDMA for Complex AI-Generated Content (AIGC) Workflows
by: Chen, June, et al.
Published: (2026) -
ALock: Asymmetric Lock Primitive for RDMA Systems
by: Baran, Amanda, et al.
Published: (2024)