Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yi, Qian, Chen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917118062100480
author Liu, Yi
Qian, Chen
author_facet Liu, Yi
Qian, Chen
contents Vector similarity search has become a critical component in AI-driven applications such as large language models (LLMs). To achieve high recall and low latency, GPUs are utilized to exploit massive parallelism for faster query processing. However, as the number of vectors continues to grow, the graph size quickly exceeds the memory capacity of a single GPU, making it infeasible to store and process the entire index on a single GPU. Recent work uses CPU-GPU architectures to keep vectors in CPU memory or SSDs, but the loading step stalls GPU computation. We present Fantasy, an efficient system that pipelines vector search and data transfer in a GPU cluster with GPUDirect Async. Fantasy overlaps computation and network communication to significantly improve search throughput for large graphs and deliver large query batch sizes.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02278
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
Liu, Yi
Qian, Chen
Distributed, Parallel, and Cluster Computing
Vector similarity search has become a critical component in AI-driven applications such as large language models (LLMs). To achieve high recall and low latency, GPUs are utilized to exploit massive parallelism for faster query processing. However, as the number of vectors continues to grow, the graph size quickly exceeds the memory capacity of a single GPU, making it infeasible to store and process the entire index on a single GPU. Recent work uses CPU-GPU architectures to keep vectors in CPU memory or SSDs, but the loading step stalls GPU computation. We present Fantasy, an efficient system that pipelines vector search and data transfer in a GPU cluster with GPUDirect Async. Fantasy overlaps computation and network communication to significantly improve search throughput for large graphs and deliver large query batch sizes.
title Fantasy: Efficient Large-scale Vector Search on GPU Clusters with GPUDirect Async
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2512.02278