Characterizing the Dilemma of Performance and Index Size in Billion-Scale Vector Search and Breaking It with Second-Tier Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Cheng, Rongxin, Peng, Yifan, Wei, Xingda, Xie, Hongrui, Chen, Rong, Shen, Sijie, Chen, Haibo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SIVF: GPU-Resident IVF Index for Streaming Vector Search
by: Zhao, Dongfang
Published: (2026)
by: Zhao, Dongfang
Published: (2026)
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024)
by: Cheng, Rongxin, et al.
Published: (2024)
Curator: Efficient Indexing for Multi-Tenant Vector Databases
by: Jin, Yicheng, et al.
Published: (2024)
by: Jin, Yicheng, et al.
Published: (2024)
PiPNN: Ultra-Scalable Graph-Based Nearest Neighbor Indexing
by: Rubel, Tobias, et al.
Published: (2026)
by: Rubel, Tobias, et al.
Published: (2026)
DIMS: Distributed Index for Similarity Search in Metric Spaces
by: Zhu, Yifan, et al.
Published: (2024)
by: Zhu, Yifan, et al.
Published: (2024)
Data Dams: A Novel Framework for Regulating and Managing Data Flow in Large-Scale Systems
by: Bouke, Mohamed Aly, et al.
Published: (2025)
by: Bouke, Mohamed Aly, et al.
Published: (2025)
Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor Search
by: Li, Zhonggen, et al.
Published: (2025)
by: Li, Zhonggen, et al.
Published: (2025)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
DiFache: Efficient and Scalable Caching on Disaggregated Memory using Decentralized Coherence
by: Zhang, Hanze, et al.
Published: (2025)
by: Zhang, Hanze, et al.
Published: (2025)
A Context-Aware Knowledge Graph Platform for Stream Processing in Industrial IoT
by: Sciarroni, Monica Marconi, et al.
Published: (2026)
by: Sciarroni, Monica Marconi, et al.
Published: (2026)
TierBase: A Workload-Driven Cost-Optimized Key-Value Store
by: Shen, Zhitao, et al.
Published: (2025)
by: Shen, Zhitao, et al.
Published: (2025)
DEX: Scalable Range Indexing on Disaggregated Memory [Extended Version]
by: Lu, Baotong, et al.
Published: (2024)
by: Lu, Baotong, et al.
Published: (2024)
SQUASH: Serverless and Distributed Quantization-based Attributed Vector Similarity Search
by: Oakley, Joe, et al.
Published: (2025)
by: Oakley, Joe, et al.
Published: (2025)
CleANN: Efficient Full Dynamism in Graph-based Approximate Nearest Neighbor Search
by: Zhang, Ziyu, et al.
Published: (2025)
by: Zhang, Ziyu, et al.
Published: (2025)
Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data Analytics
by: Yuan, Yichao, et al.
Published: (2025)
by: Yuan, Yichao, et al.
Published: (2025)
KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud Provider
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Data Caching for Enterprise-Grade Petabyte-Scale OLAP
by: Tang, Chunxu, et al.
Published: (2024)
by: Tang, Chunxu, et al.
Published: (2024)
Learning from the Past: Adaptive Parallelism Tuning for Stream Processing Systems
by: Han, Yuxing, et al.
Published: (2025)
by: Han, Yuxing, et al.
Published: (2025)
Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Evaluating the Impact Of Spatial Features Of Mobility Data and Index Choice On Database Performance
by: Rese, Tim C., et al.
Published: (2025)
by: Rese, Tim C., et al.
Published: (2025)
Distributed Indexing Schemes for k-Dominant Skyline Analytics on Uncertain Edge-IoT Data
by: Lai, Chuan-Chi, et al.
Published: (2023)
by: Lai, Chuan-Chi, et al.
Published: (2023)
A Pragmatic Approach to Learned Indexing in RocksDB: Targeted Optimizations with Minimal System Modification
by: Vashisth, Shubham, et al.
Published: (2026)
by: Vashisth, Shubham, et al.
Published: (2026)
CIDER: Boosting Memory-Disaggregated Key-Value Stores with Pessimistic Synchronization
by: Du, Yuxuan, et al.
Published: (2026)
by: Du, Yuxuan, et al.
Published: (2026)
Next Generation Cloud-native In-Memory Stores: From Redis to Valkey and Beyond
by: Rosensch"old, Carl-Johan Fauvelle Munck af, et al.
Published: (2025)
by: Rosensch"old, Carl-Johan Fauvelle Munck af, et al.
Published: (2025)
TeraHAC: Hierarchical Agglomerative Clustering of Trillion-Edge Graphs
by: Dhulipala, Laxman, et al.
Published: (2023)
by: Dhulipala, Laxman, et al.
Published: (2023)
Parallel R-tree-based Spatial Query Processing on a Commercial Processing-in-Memory System
by: Jannat, Tasmia, et al.
Published: (2026)
by: Jannat, Tasmia, et al.
Published: (2026)
Delta Tensor: Efficient Vector and Tensor Storage in Delta Lake
by: Bao, Zhiwei, et al.
Published: (2024)
by: Bao, Zhiwei, et al.
Published: (2024)
Efficient Batch Search Algorithm for B+ Tree Index Structures with Level-Wise Traversal on FPGAs
by: Tzschoppe, Max, et al.
Published: (2026)
by: Tzschoppe, Max, et al.
Published: (2026)
StreamShield: A Production-Proven Resiliency Solution for Apache Flink at ByteDance
by: Fang, Yong, et al.
Published: (2026)
by: Fang, Yong, et al.
Published: (2026)
Exploring Novel Data Storage Approaches for Large-Scale Numerical Weather Prediction
by: Gil, Nicolau Manubens
Published: (2026)
by: Gil, Nicolau Manubens
Published: (2026)
PolarStore: High-Performance Data Compression for Large-Scale Cloud-Native Databases
by: Hu, Qingda, et al.
Published: (2025)
by: Hu, Qingda, et al.
Published: (2025)
LatentBox: Storing AI-Generated Images at Scale via a Latent-First Design
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
ACGraph: An Efficient Asynchronous Out-of-Core Graph Processing Framework
by: Chen, Dechuang, et al.
Published: (2025)
by: Chen, Dechuang, et al.
Published: (2025)
Do GPUs Really Need New Tabular File Formats?
by: Luo, Jigao, et al.
Published: (2026)
by: Luo, Jigao, et al.
Published: (2026)
AgileDART: An Agile and Scalable Edge Stream Processing Engine
by: Ching, Cheng-Wei, et al.
Published: (2024)
by: Ching, Cheng-Wei, et al.
Published: (2024)
Exact Nearest-Neighbor Search on Energy-Efficient FPGA Devices
by: Dazzi, Patrizio, et al.
Published: (2025)
by: Dazzi, Patrizio, et al.
Published: (2025)
Did we miss P In CAP? Partial Progress Conjecture under Asynchrony
by: Chen, Junchao, et al.
Published: (2024)
by: Chen, Junchao, et al.
Published: (2024)
Taurus Database: How to be Fast, Available, and Frugal in the Cloud
by: Depoutovitch, Alex, et al.
Published: (2024)
by: Depoutovitch, Alex, et al.
Published: (2024)
SFVInt: Simple, Fast and Generic Variable-Length Integer Decoding using Bit Manipulation Instructions
by: Liao, Gang, et al.
Published: (2024)
by: Liao, Gang, et al.
Published: (2024)
Lion: Minimizing Distributed Transactions through Adaptive Replica Provision (Extended Version)
by: Zheng, Qiushi, et al.
Published: (2024)
by: Zheng, Qiushi, et al.
Published: (2024)
Similar Items
-
SIVF: GPU-Resident IVF Index for Streaming Vector Search
by: Zhao, Dongfang
Published: (2026) -
KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
by: Cheng, Rongxin, et al.
Published: (2024) -
Curator: Efficient Indexing for Multi-Tenant Vector Databases
by: Jin, Yicheng, et al.
Published: (2024) -
PiPNN: Ultra-Scalable Graph-Based Nearest Neighbor Indexing
by: Rubel, Tobias, et al.
Published: (2026) -
DIMS: Distributed Index for Similarity Search in Metric Spaces
by: Zhu, Yifan, et al.
Published: (2024)