Efficient Graph-Based Approximate Nearest Neighbor Search Achieving: Low Latency Without Throughput Loss
Fuente:
arXiv
Guardado en:
| Autores principales: | Luo, Jingjia, Zhang, Mingxing, Chen, Kang, Liao, Xia, Shan, Yingdi, Jiang, Jinlei, Wu, Yongwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
por: Chen, Shaoyuan, et al.
Publicado: (2024)
por: Chen, Shaoyuan, et al.
Publicado: (2024)
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
por: Kim, Sukjin, et al.
Publicado: (2025)
por: Kim, Sukjin, et al.
Publicado: (2025)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
por: Li, Zhonggen, et al.
Publicado: (2025)
por: Li, Zhonggen, et al.
Publicado: (2025)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
por: V., Karthik, et al.
Publicado: (2024)
por: V., Karthik, et al.
Publicado: (2024)
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
TrEnv-X: Transparently Share Serverless Execution Environments Across Different Functions and Nodes
por: Huang, Jialiang, et al.
Publicado: (2025)
por: Huang, Jialiang, et al.
Publicado: (2025)
PECANN: Parallel Efficient Clustering with Graph-Based Approximate Nearest Neighbor Search
por: Yu, Shangdi, et al.
Publicado: (2023)
por: Yu, Shangdi, et al.
Publicado: (2023)
Advancing RT Core-Accelerated Fixed-Radius Nearest Neighbor Search
por: Meneses, Enzo, et al.
Publicado: (2026)
por: Meneses, Enzo, et al.
Publicado: (2026)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
por: Dazzi, Patrizio, et al.
Publicado: (2025)
por: Dazzi, Patrizio, et al.
Publicado: (2025)
CleANN: Efficient Full Dynamism in Graph-based Approximate Nearest Neighbor Search
por: Zhang, Ziyu, et al.
Publicado: (2025)
por: Zhang, Ziyu, et al.
Publicado: (2025)
BBCA-CHAIN: Low Latency, High Throughput BFT Consensus on a DAG
por: Malkhi, Dahlia, et al.
Publicado: (2023)
por: Malkhi, Dahlia, et al.
Publicado: (2023)
Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic Workloads
por: Hidayetoglu, Mert, et al.
Publicado: (2025)
por: Hidayetoglu, Mert, et al.
Publicado: (2025)
Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
por: Qin, Ruoyu, et al.
Publicado: (2025)
por: Qin, Ruoyu, et al.
Publicado: (2025)
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching
por: Ruan, Chaoyi, et al.
Publicado: (2025)
por: Ruan, Chaoyi, et al.
Publicado: (2025)
Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced Throughput
por: Dai, Xiaohai, et al.
Publicado: (2025)
por: Dai, Xiaohai, et al.
Publicado: (2025)
Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference
por: Yu, Minchen, et al.
Publicado: (2023)
por: Yu, Minchen, et al.
Publicado: (2023)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
por: Mandarapu, Durga, et al.
Publicado: (2023)
por: Mandarapu, Durga, et al.
Publicado: (2023)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
por: Zou, Cheng, et al.
Publicado: (2026)
por: Zou, Cheng, et al.
Publicado: (2026)
PiPNN: Ultra-Scalable Graph-Based Nearest Neighbor Indexing
por: Rubel, Tobias, et al.
Publicado: (2026)
por: Rubel, Tobias, et al.
Publicado: (2026)
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
por: Wang, Jun, et al.
Publicado: (2025)
por: Wang, Jun, et al.
Publicado: (2025)
On the Effectiveness of Graph Reordering for Accelerating Approximate Nearest Neighbor Search on GPU
por: Oguri, Yutaro, et al.
Publicado: (2025)
por: Oguri, Yutaro, et al.
Publicado: (2025)
Low-Latency Federated Fine-Tuning for Large Language Models Over Wireless Networks
por: Pang, Zhiwen, et al.
Publicado: (2026)
por: Pang, Zhiwen, et al.
Publicado: (2026)
Fast Iterative Graph Computing with Updated Neighbor States
por: Zhou, Yijie, et al.
Publicado: (2024)
por: Zhou, Yijie, et al.
Publicado: (2024)
Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference
por: Sun, Xun, et al.
Publicado: (2026)
por: Sun, Xun, et al.
Publicado: (2026)
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter
por: Qin, Ruoyu, et al.
Publicado: (2026)
por: Qin, Ruoyu, et al.
Publicado: (2026)
CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs
por: Ootomo, Hiroyuki, et al.
Publicado: (2023)
por: Ootomo, Hiroyuki, et al.
Publicado: (2023)
SOLANET: Distributed Neighbor Graph Construction on GPU-Accelerated Systems
por: Iwabuchi, Keita, et al.
Publicado: (2026)
por: Iwabuchi, Keita, et al.
Publicado: (2026)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
por: Agrawal, Amey, et al.
Publicado: (2024)
por: Agrawal, Amey, et al.
Publicado: (2024)
Revisiting Speculative Leaderless Protocols for Low-Latency BFT Replication
por: Qian, Daniel, et al.
Publicado: (2026)
por: Qian, Daniel, et al.
Publicado: (2026)
Low Latency, High Bandwidth Streaming of Experimental Data with EJFAT
por: Baldin, Ilya, et al.
Publicado: (2025)
por: Baldin, Ilya, et al.
Publicado: (2025)
NAVIS: Concurrent Search and Update with Low Position-Seeking Overhead in On-SSD Graph-Based Vector Search
por: Song, Jaeyong, et al.
Publicado: (2026)
por: Song, Jaeyong, et al.
Publicado: (2026)
Dumbo-NG: Fast Asynchronous BFT Consensus with Throughput-Oblivious Latency
por: Gao, Yingzi, et al.
Publicado: (2022)
por: Gao, Yingzi, et al.
Publicado: (2022)
TENT: A Declarative Slice Spraying Engine for Performant and Resilient Data Movement in Disaggregated LLM Serving
por: Ren, Feng, et al.
Publicado: (2026)
por: Ren, Feng, et al.
Publicado: (2026)
Low-Latency Layer-Aware Proactive and Passive Container Migration in Meta Computing
por: Liu, Mengjie, et al.
Publicado: (2024)
por: Liu, Mengjie, et al.
Publicado: (2024)
Chasing the Speed of Light: Low-Latency Planetary-Scale Adaptive Byzantine Consensus
por: Berger, Christian, et al.
Publicado: (2023)
por: Berger, Christian, et al.
Publicado: (2023)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
por: Song, Jingwei, et al.
Publicado: (2025)
por: Song, Jingwei, et al.
Publicado: (2025)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
por: Dai, Yinwei, et al.
Publicado: (2023)
por: Dai, Yinwei, et al.
Publicado: (2023)
Scene-Aware Latency Estimation for Microservices via Multi-Scale Graph Fusion
por: Sun, Zhichao, et al.
Publicado: (2026)
por: Sun, Zhichao, et al.
Publicado: (2026)
OMEGA: A Low-Latency GNN Serving System for Large Graphs
por: Kim, Geon-Woo, et al.
Publicado: (2025)
por: Kim, Geon-Woo, et al.
Publicado: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
por: Wang, Wenfeng, et al.
Publicado: (2026)
por: Wang, Wenfeng, et al.
Publicado: (2026)
Ejemplares similares
-
Efficient Heterogeneous Large Language Model Decoding with Model-Attention Disaggregation
por: Chen, Shaoyuan, et al.
Publicado: (2024) -
PathWeaver: A High-Throughput Multi-GPU System for Graph-Based Approximate Nearest Neighbor Search
por: Kim, Sukjin, et al.
Publicado: (2025) -
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
por: Li, Zhonggen, et al.
Publicado: (2025) -
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
por: V., Karthik, et al.
Publicado: (2024) -
GRNND: A GPU-Parallel Relative NN-Descent Algorithm for Efficient Approximate Nearest Neighbor Graph Construction
por: Li, Xiang, et al.
Publicado: (2025)