HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage
Fuente:
arXiv
Saved in:
| Main Authors: | Rong, Haidong, Yao, Jiashu, Langer, Matthias, Liu, Shijie, Fan, Li, Wang, Dongxin, He, Jia, Chen, Jinglin, Rang, Jiaheng, Qian, Julian, Xu, Mengyao, Yu, Fan, Lee, Minseok, Wang, Zehuan, Oldridge, Even |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
by: Kirdey, Stanislav, et al.
Published: (2026)
by: Kirdey, Stanislav, et al.
Published: (2026)
Beyond the Geometric Curse: High-Dimensional N-Gram Hashing for Dense Retrieval
by: Sharma, Sangeet
Published: (2026)
by: Sharma, Sangeet
Published: (2026)
When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
by: Ghaffari, Shervin, et al.
Published: (2025)
by: Ghaffari, Shervin, et al.
Published: (2025)
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
by: Xu, Yichun, et al.
Published: (2026)
by: Xu, Yichun, et al.
Published: (2026)
QVCache: A Query-Aware Vector Cache
by: Göçer, Anıl Eren, et al.
Published: (2026)
by: Göçer, Anıl Eren, et al.
Published: (2026)
Fine-Grained Embedding Dimension Optimization During Training for Recommender Systems
by: Luo, Qinyi, et al.
Published: (2024)
by: Luo, Qinyi, et al.
Published: (2024)
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
AWARE: Evaluating PriorityFresh Caching for Offline Emergency Warning Systems
by: Melvin, Charles, et al.
Published: (2025)
by: Melvin, Charles, et al.
Published: (2025)
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Modular GPU Programming with Typed Perspectives
by: Bansal, Manya, et al.
Published: (2025)
by: Bansal, Manya, et al.
Published: (2025)
The Deterministic Horizon: Impossibility Results as Design Specifications for Trustworthy AI Systems
by: Guo, Dongxin
Published: (2026)
by: Guo, Dongxin
Published: (2026)
Categorical Calculus and Algebra for Multi-Model Data
by: Lu, Jiaheng
Published: (2026)
by: Lu, Jiaheng
Published: (2026)
Llama-Embed-Nemotron-8B: A Universal Text Embedding Model for Multilingual and Cross-Lingual Tasks
by: Babakhin, Yauhen, et al.
Published: (2025)
by: Babakhin, Yauhen, et al.
Published: (2025)
When Hard Negatives Hurt: Bridging the Generative-Discriminative Gap in Hard Negative Synthesis for Retrieval
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Semantic Caching for OLAP via LLM-Based Query Canonicalization (Extended Version)
by: Bindschaedler, Laurent
Published: (2026)
by: Bindschaedler, Laurent
Published: (2026)
GPU-friendly Stroke Expansion
by: Levien, Raph, et al.
Published: (2024)
by: Levien, Raph, et al.
Published: (2024)
Relevance Filtering for Embedding-based Retrieval
by: Rossi, Nicholas, et al.
Published: (2024)
by: Rossi, Nicholas, et al.
Published: (2024)
Timehash: Hierarchical Time Indexing for Efficient Business Hours Search
by: Kim, Jinoh, et al.
Published: (2026)
by: Kim, Jinoh, et al.
Published: (2026)
Enhancing Relevance of Embedding-based Retrieval at Walmart
by: Lin, Juexin, et al.
Published: (2024)
by: Lin, Juexin, et al.
Published: (2024)
Enabling full-speed random access to the entire memory on the A100 GPU
by: Walker, Alden
Published: (2024)
by: Walker, Alden
Published: (2024)
Specular Polynomials
by: Fan, Zhimin, et al.
Published: (2024)
by: Fan, Zhimin, et al.
Published: (2024)
Cluster-based Graph Collaborative Filtering
by: Liu, Fan, et al.
Published: (2024)
by: Liu, Fan, et al.
Published: (2024)
flexvec: SQL Vector Retrieval with Programmatic Embedding Modulation
by: Delmas, Damian
Published: (2026)
by: Delmas, Damian
Published: (2026)
Make LLM Inference Affordable to Everyone: Augmenting GPU Memory with NDP-DIMM
by: Liu, Lian, et al.
Published: (2025)
by: Liu, Lian, et al.
Published: (2025)
I$^3$-MRec: Invariant Learning with Information Bottleneck for Incomplete Modality Recommendation
by: Chen, Huilin, et al.
Published: (2025)
by: Chen, Huilin, et al.
Published: (2025)
High-Frequency-aware Hierarchical Contrastive Selective Coding for Representation Learning on Text-attributed Graphs
by: Zhang, Peiyan, et al.
Published: (2024)
by: Zhang, Peiyan, et al.
Published: (2024)
A PLMs based protein retrieval framework
by: Wu, Yuxuan, et al.
Published: (2024)
by: Wu, Yuxuan, et al.
Published: (2024)
CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry
by: Ye, Xiaopeng, et al.
Published: (2025)
by: Ye, Xiaopeng, et al.
Published: (2025)
CSMF: Cascaded Selective Mask Fine-Tuning for Multi-Objective Embedding-Based Retrieval
by: Deng, Hao, et al.
Published: (2025)
by: Deng, Hao, et al.
Published: (2025)
CardioEmbed: Domain-Specialized Text Embeddings for Clinical Cardiology
by: Young, Richard J., et al.
Published: (2025)
by: Young, Richard J., et al.
Published: (2025)
DAInfer+: Neurosymbolic Inference of API Specifications from Documentation via Embedding Models
by: Masoudian, Maryam, et al.
Published: (2026)
by: Masoudian, Maryam, et al.
Published: (2026)
Understanding Before Recommendation: Semantic Aspect-Aware Review Exploitation via Large Language Models
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
KamerRaad: Enhancing Information Retrieval in Belgian National Politics through Hierarchical Summarization and Conversational Interfaces
by: Rogiers, Alexander, et al.
Published: (2024)
by: Rogiers, Alexander, et al.
Published: (2024)
Bandicoot: A Templated C++ Library for GPU Linear Algebra
by: Curtin, Ryan R., et al.
Published: (2025)
by: Curtin, Ryan R., et al.
Published: (2025)
Content-Aware Tweet Location Inference using Quadtree Spatial Partitioning and Jaccard-Cosine Word Embedding
by: Ajao, Oluwaseun, et al.
Published: (2024)
by: Ajao, Oluwaseun, et al.
Published: (2024)
CLM: Removing the GPU Memory Barrier for 3D Gaussian Splatting
by: Zhao, Hexu, et al.
Published: (2025)
by: Zhao, Hexu, et al.
Published: (2025)
Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Session Context Embedding for Intent Understanding in Product Search
by: Mehrdad, Navid, et al.
Published: (2024)
by: Mehrdad, Navid, et al.
Published: (2024)
Similar Items
-
Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
by: Kirdey, Stanislav, et al.
Published: (2026) -
Beyond the Geometric Curse: High-Dimensional N-Gram Hashing for Dense Retrieval
by: Sharma, Sangeet
Published: (2026) -
When to Retrieve During Reasoning: Adaptive Retrieval for Large Reasoning Models
by: Guo, Dongxin, et al.
Published: (2026) -
An Ensemble Embedding Approach for Improving Semantic Caching Performance in LLM-based Systems
by: Ghaffari, Shervin, et al.
Published: (2025) -
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization
by: Guo, Dongxin, et al.
Published: (2026)