Incremental IVF Index Maintenance for Streaming Vector Search
Fuente:
arXiv
Guardado en:
| Autores principales: | Mohoney, Jason, Pacaci, Anil, Chowdhury, Shihabur Rahman, Minhas, Umar Farooq, Pound, Jeffery, Renggli, Cedric, Reyhani, Nima, Ilyas, Ihab F., Rekatsinas, Theodoros, Venkataraman, Shivaram |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Quake: Adaptive Indexing for Vector Search
por: Mohoney, Jason, et al.
Publicado: (2025)
por: Mohoney, Jason, et al.
Publicado: (2025)
Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
por: Renggli, Cedric, et al.
Publicado: (2025)
por: Renggli, Cedric, et al.
Publicado: (2025)
MicroNN: An On-device Disk-resident Updatable Vector Database
por: Pound, Jeffrey, et al.
Publicado: (2025)
por: Pound, Jeffrey, et al.
Publicado: (2025)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
por: Waleffe, Roger, et al.
Publicado: (2025)
por: Waleffe, Roger, et al.
Publicado: (2025)
TSDS: Data Selection for Task-Specific Model Finetuning
por: Liu, Zifan, et al.
Publicado: (2024)
por: Liu, Zifan, et al.
Publicado: (2024)
SIVF: GPU-Resident IVF Index for Streaming Vector Search
por: Zhao, Dongfang
Publicado: (2026)
por: Zhao, Dongfang
Publicado: (2026)
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
Eva: Cost-Efficient Cloud-Based Cluster Scheduling
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
por: Chang, Tzu-Tao, et al.
Publicado: (2025)
Scaling Inference-Efficient Language Models
por: Bian, Song, et al.
Publicado: (2025)
por: Bian, Song, et al.
Publicado: (2025)
Decoding Speculative Decoding
por: Yan, Minghao, et al.
Publicado: (2024)
por: Yan, Minghao, et al.
Publicado: (2024)
PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices
por: Yan, Minghao, et al.
Publicado: (2023)
por: Yan, Minghao, et al.
Publicado: (2023)
Analyzing the Evolution and Maintenance of Quantum Software Repositories
por: Upadhyay, Krishna, et al.
Publicado: (2025)
por: Upadhyay, Krishna, et al.
Publicado: (2025)
IVF-TQ: Calibration-Free Streaming Vector Search via a Codebook-Free Residual Layer
por: Sharma, Tarun
Publicado: (2026)
por: Sharma, Tarun
Publicado: (2026)
From FASTER to F2: Evolving Concurrent Key-Value Store Designs for Large Skewed Workloads
por: Kanellis, Konstantinos, et al.
Publicado: (2023)
por: Kanellis, Konstantinos, et al.
Publicado: (2023)
Scaling Laws Meet Model Architecture: Toward Inference-Efficient LLMs
por: Bian, Song, et al.
Publicado: (2025)
por: Bian, Song, et al.
Publicado: (2025)
GraphSnapShot: Caching Local Structure for Fast Graph Learning
por: Liu, Dong, et al.
Publicado: (2024)
por: Liu, Dong, et al.
Publicado: (2024)
ARMS: Adaptive and Robust Memory Tiering System
por: Yadalam, Sujay, et al.
Publicado: (2025)
por: Yadalam, Sujay, et al.
Publicado: (2025)
TUNA: Tuning Unstable and Noisy Cloud Applications
por: Freischuetz, Johannes, et al.
Publicado: (2025)
por: Freischuetz, Johannes, et al.
Publicado: (2025)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Real Time Password Strength Analysis on a Web Application Using Multiple Machine Learning Approaches
por: Umar Farooq
Publicado: (2020)
por: Umar Farooq
Publicado: (2020)
Recent Increments in Incremental View Maintenance
por: Olteanu, Dan
Publicado: (2024)
por: Olteanu, Dan
Publicado: (2024)
Tesserae: Scalable Placement Policies for Deep Learning Workloads
por: Bian, Song, et al.
Publicado: (2025)
por: Bian, Song, et al.
Publicado: (2025)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
por: Conia, Simone, et al.
Publicado: (2024)
por: Conia, Simone, et al.
Publicado: (2024)
Exploring Distributed Vector Databases Performance on HPC Platforms: A Study with Qdrant
por: Ockerman, Seth, et al.
Publicado: (2025)
por: Ockerman, Seth, et al.
Publicado: (2025)
E-Governance and the Future of Public Service in the Digital Age
por: Muhammad Umar Farooq
Publicado: (2025)
por: Muhammad Umar Farooq
Publicado: (2025)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
por: Farooq, Muhammad Umar
Publicado: (2026)
por: Farooq, Muhammad Umar
Publicado: (2026)
Depicting Psychological Trauma: A study of Stream of Consciousness in Atiq Rahimi's A Thousand Rooms of Dreams and Fear
por: Saneen Iraj, et al.
Publicado: (2025)
por: Saneen Iraj, et al.
Publicado: (2025)
What Limits Agentic Systems Efficiency?
por: Bian, Song, et al.
Publicado: (2025)
por: Bian, Song, et al.
Publicado: (2025)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
por: Kanellis, Konstantinos, et al.
Publicado: (2025)
por: Kanellis, Konstantinos, et al.
Publicado: (2025)
PLoRA: Efficient LoRA Hyperparameter Tuning for Large Models
por: Yan, Minghao, et al.
Publicado: (2025)
por: Yan, Minghao, et al.
Publicado: (2025)
ConvKGYarn: Spinning Configurable and Scalable Conversational Knowledge Graph QA datasets with Large Language Models
por: Pradeep, Ronak, et al.
Publicado: (2024)
por: Pradeep, Ronak, et al.
Publicado: (2024)
Incremental Maintenance of DatalogMTL Materialisations
por: Zhao, Kaiyue, et al.
Publicado: (2025)
por: Zhao, Kaiyue, et al.
Publicado: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
por: Jain, Rutwik, et al.
Publicado: (2024)
por: Jain, Rutwik, et al.
Publicado: (2024)
Wattchmen: Watching the Wattchers -- High Fidelity, Flexible GPU Energy Modeling
por: Tran, Brandon, et al.
Publicado: (2026)
por: Tran, Brandon, et al.
Publicado: (2026)
Teaching Quantum Computing through Lab-Integrated Learning: Bridging Conceptual and Computational Understanding
por: Farooq, Umar, et al.
Publicado: (2025)
por: Farooq, Umar, et al.
Publicado: (2025)
Understanding and Detecting Platform-Specific Violations in Android Auto Apps
por: Fakorede, Moshood, et al.
Publicado: (2025)
por: Fakorede, Moshood, et al.
Publicado: (2025)
Quantile connectedness of artificial intelligence tokens with the energy sector
por: Farooq Malik, et al.
Publicado: (2024)
por: Farooq Malik, et al.
Publicado: (2024)
Exploring the Impact of Environmental, Social, and Governance ( ESG ) Performance on Firm Efficiency: The Mediating Role of Environmental R&D Investment in BRICS Economies
por: Umar Farooq, et al.
Publicado: (2025)
por: Umar Farooq, et al.
Publicado: (2025)
Fertility Treatment at Sneh IVF Hospital | Best IVF Center in Ahmedabad
por: Center, Sneh ivf
Publicado: (2026)
por: Center, Sneh ivf
Publicado: (2026)
Ejemplares similares
-
Quake: Adaptive Indexing for Vector Search
por: Mohoney, Jason, et al.
Publicado: (2025) -
Fundamental Challenges in Evaluating Text2SQL Solutions and Detecting Their Limitations
por: Renggli, Cedric, et al.
Publicado: (2025) -
MicroNN: An On-device Disk-resident Updatable Vector Database
por: Pound, Jeffrey, et al.
Publicado: (2025) -
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
por: Waleffe, Roger, et al.
Publicado: (2025) -
TSDS: Data Selection for Task-Specific Model Finetuning
por: Liu, Zifan, et al.
Publicado: (2024)