FusionANNS: An Efficient CPU/GPU Cooperative Processing Architecture for Billion-scale Approximate Nearest Neighbor Search
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Bing, Liu, Haikun, Tang, Yuhang, Xiao, Shihai, Duan, Zhuohui, Liao, Xiaofei, Zhang, Xuecang, Zhu, Junhua, Zhang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers
von: Liu, Mugeng, et al.
Veröffentlicht: (2025)
von: Liu, Mugeng, et al.
Veröffentlicht: (2025)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
von: V., Karthik, et al.
Veröffentlicht: (2024)
von: V., Karthik, et al.
Veröffentlicht: (2024)
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
von: Yang, Yiwei, et al.
Veröffentlicht: (2026)
von: Yang, Yiwei, et al.
Veröffentlicht: (2026)
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)
RUISA Operational Ecosystem Architecture
von: AL Mohtar, Mouayad
Veröffentlicht: (2026)
von: AL Mohtar, Mouayad
Veröffentlicht: (2026)
UrgenGo: Urgency-Aware Transparent GPU Kernel Launching for Autonomous Driving
von: Zhu, Hanqi, et al.
Veröffentlicht: (2025)
von: Zhu, Hanqi, et al.
Veröffentlicht: (2025)
Optimizing Task Scheduling in Heterogeneous Computing Environments: A Comparative Analysis of CPU, GPU, and ASIC Platforms Using E2C Simulator
von: Mohammadjafari, Ali, et al.
Veröffentlicht: (2024)
von: Mohammadjafari, Ali, et al.
Veröffentlicht: (2024)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
von: Siavashi, Mohammad, et al.
Veröffentlicht: (2026)
Exploiting Dependency and Parallelism: Real-Time Scheduling and Analysis for GPU Tasks
von: Zhang, Yuanhai, et al.
Veröffentlicht: (2026)
von: Zhang, Yuanhai, et al.
Veröffentlicht: (2026)
Flare: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus Scale
von: Cui, Weihao, et al.
Veröffentlicht: (2025)
von: Cui, Weihao, et al.
Veröffentlicht: (2025)
UpANNS: Enhancing Billion-Scale ANNS Efficiency with Real-World PIM Architecture
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
von: Chen, Sitian, et al.
Veröffentlicht: (2024)
The K$K$‐prize‐collecting coverage problem by aligned disks
von: Hao Zhang, et al.
Veröffentlicht: (2025)
von: Hao Zhang, et al.
Veröffentlicht: (2025)
Towards Fully-fledged GPU Multitasking via Proactive Memory Scheduling
von: Shen, Weihang, et al.
Veröffentlicht: (2025)
von: Shen, Weihang, et al.
Veröffentlicht: (2025)
RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
von: Yu, Weiping, et al.
Veröffentlicht: (2025)
Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization
von: Xu, Yuhang, et al.
Veröffentlicht: (2025)
von: Xu, Yuhang, et al.
Veröffentlicht: (2025)
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
von: Li, Xiaoya, et al.
Veröffentlicht: (2025)
Trustworthy and Controllable Professional Knowledge Utilization in Large Language Models with TEE-GPU Execution
von: Cai, Yifeng, et al.
Veröffentlicht: (2025)
von: Cai, Yifeng, et al.
Veröffentlicht: (2025)
CPU-Limits kill Performance: Time to rethink Resource Control
von: Shetty, Chirag, et al.
Veröffentlicht: (2025)
von: Shetty, Chirag, et al.
Veröffentlicht: (2025)
On the Effectiveness of Graph Reordering for Accelerating Approximate Nearest Neighbor Search on GPU
von: Oguri, Yutaro, et al.
Veröffentlicht: (2025)
von: Oguri, Yutaro, et al.
Veröffentlicht: (2025)
Co-Designing Graph-based Approximate Nearest Neighbor Search at Billion Scale for Processing-in-Memory
von: Chen, Sitian, et al.
Veröffentlicht: (2026)
von: Chen, Sitian, et al.
Veröffentlicht: (2026)
Microsecond-scale Dynamic Validation of Idempotency for GPU Kernels
von: Han, Mingcong, et al.
Veröffentlicht: (2024)
von: Han, Mingcong, et al.
Veröffentlicht: (2024)
Breaking the Storage-Compute Bottleneck in Billion-Scale ANNS: A GPU-Driven Asynchronous I/O Framework
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
Existence of Approximately Macroscopically Unique States
von: Lin, Huaxin
Veröffentlicht: (2024)
von: Lin, Huaxin
Veröffentlicht: (2024)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
von: Tofigh, Mani, et al.
Veröffentlicht: (2025)
Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices
von: Liang, Yu, et al.
Veröffentlicht: (2025)
von: Liang, Yu, et al.
Veröffentlicht: (2025)
Approximately Macroscopically Unique States and Quantum Mechanics
von: Lin, Huaxin, et al.
Veröffentlicht: (2025)
von: Lin, Huaxin, et al.
Veröffentlicht: (2025)
Energy-Aware CPU Orchestration in O-RAN: A dApp-Driven Lightweight Approach
von: Crespo, Francisco, et al.
Veröffentlicht: (2025)
von: Crespo, Francisco, et al.
Veröffentlicht: (2025)
RTOS Architectures that Solve the Diminishing Bandwidth Problem
von: Arakji, Mazen
Veröffentlicht: (2025)
von: Arakji, Mazen
Veröffentlicht: (2025)
Approximation by elements of finite spectra for C* Algebras of higher real rank
von: Sarkar, Aranya
Veröffentlicht: (2025)
von: Sarkar, Aranya
Veröffentlicht: (2025)
Circulant quantum channels and its applications
von: Xie, Bing, et al.
Veröffentlicht: (2026)
von: Xie, Bing, et al.
Veröffentlicht: (2026)
Exploiting Application-to-Architecture Dependencies for Designing Scalable OS
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
Complete Metric Approximation Property For Mixed $q$-Deformed Araki-Woods Factors
von: Bikram, Panchugopal, et al.
Veröffentlicht: (2024)
von: Bikram, Panchugopal, et al.
Veröffentlicht: (2024)
SSV: Sparse Speculative Verification for Efficient LLM Inference
von: Wang, Zhibin, et al.
Veröffentlicht: (2026)
von: Wang, Zhibin, et al.
Veröffentlicht: (2026)
Generalised diagonal dimension and applications to large-scale geometry
von: Kitsios, Christos
Veröffentlicht: (2026)
von: Kitsios, Christos
Veröffentlicht: (2026)
On the Completely Positive Approximation Property for Non-Unital Operator Systems and the Boundary Condition for the Zero Map
von: Kim, Se-Jin
Veröffentlicht: (2024)
von: Kim, Se-Jin
Veröffentlicht: (2024)
PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
von: Song, Yixin, et al.
Veröffentlicht: (2023)
von: Song, Yixin, et al.
Veröffentlicht: (2023)
From Stable Rank One to Real Rank Zero: A Note on Tracial Approximate Oscillation Zero
von: Fu, Xuanlong
Veröffentlicht: (2025)
von: Fu, Xuanlong
Veröffentlicht: (2025)
Characterizing Network Requirements for GPU API Remoting in AI Applications
von: Wang, Tianxia, et al.
Veröffentlicht: (2024)
von: Wang, Tianxia, et al.
Veröffentlicht: (2024)
The choice of financing mechanism in reward‐based crowdfunding: A perspective on information disclosure
von: Mingzhao Tang, et al.
Veröffentlicht: (2025)
von: Mingzhao Tang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WebANNS: Fast and Efficient Approximate Nearest Neighbor Search in Web Browsers
von: Liu, Mugeng, et al.
Veröffentlicht: (2025) -
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
von: V., Karthik, et al.
Veröffentlicht: (2024) -
GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion
von: Yang, Yiwei, et al.
Veröffentlicht: (2026) -
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026) -
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
von: Kamahori, Keisuke, et al.
Veröffentlicht: (2024)