HACache: Leveraging Read Performance with Cache in a Heterogeneous Array
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Jialin, Shi, Liang, Yu, Dingcui |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
por: Yu, Dingcui, et al.
Publicado: (2025)
por: Yu, Dingcui, et al.
Publicado: (2025)
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
por: Qiu, Shi, et al.
Publicado: (2026)
por: Qiu, Shi, et al.
Publicado: (2026)
Cache is King: Smart Page Eviction with eBPF
por: Zussman, Tal, et al.
Publicado: (2025)
por: Zussman, Tal, et al.
Publicado: (2025)
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
por: Wu, Fangnuo, et al.
Publicado: (2025)
por: Wu, Fangnuo, et al.
Publicado: (2025)
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
por: Kanellopoulos, Konstantinos, et al.
Publicado: (2023)
por: Kanellopoulos, Konstantinos, et al.
Publicado: (2023)
DynamicAdaptiveClimb: Adaptive Cache Replacement with Dynamic Resizing
por: Berend, Daniel, et al.
Publicado: (2025)
por: Berend, Daniel, et al.
Publicado: (2025)
Oneiros: KV Cache Optimization through Parameter Remapping for Multi-tenant LLM Serving
por: Li, Ruihao, et al.
Publicado: (2025)
por: Li, Ruihao, et al.
Publicado: (2025)
DFUSE: Strongly Consistent Write-Back Kernel Caching for Distributed Userspace File Systems
por: Li, Haoyu, et al.
Publicado: (2025)
por: Li, Haoyu, et al.
Publicado: (2025)
2DIO: A Cache-Accurate Storage Microbenchmark
por: Wang, Yirong, et al.
Publicado: (2026)
por: Wang, Yirong, et al.
Publicado: (2026)
Concurrency Testing in the Linux Kernel via eBPF
por: Xu, Jiacheng, et al.
Publicado: (2025)
por: Xu, Jiacheng, et al.
Publicado: (2025)
Leveraging OS-Level Primitives for Robotic Action Management
por: Zheng, Wenxin, et al.
Publicado: (2025)
por: Zheng, Wenxin, et al.
Publicado: (2025)
RTP-LLM: High-Performance Alibaba LLM Inference Engine
por: Tan, Boyu, et al.
Publicado: (2026)
por: Tan, Boyu, et al.
Publicado: (2026)
Optimizing Tree-structure Indexes for CXL-based Heterogeneous Memory with SINLK
por: Zhao, Haoru, et al.
Publicado: (2025)
por: Zhao, Haoru, et al.
Publicado: (2025)
Principled Performance Tunability in Operating System Kernels
por: Chen, Zhongjie, et al.
Publicado: (2025)
por: Chen, Zhongjie, et al.
Publicado: (2025)
Dissecting CXL Memory Performance at Scale: Analysis, Modeling, and Optimization
por: Liu, Jinshu, et al.
Publicado: (2024)
por: Liu, Jinshu, et al.
Publicado: (2024)
Holistic Heterogeneous Scheduling for Autonomous Applications using Fine-grained, Multi-XPU Abstraction
por: Han, Mingcong, et al.
Publicado: (2025)
por: Han, Mingcong, et al.
Publicado: (2025)
E-Mapper: Energy-Efficient Resource Allocation for Traditional Operating Systems on Heterogeneous Processors
por: Smejkal, Till, et al.
Publicado: (2024)
por: Smejkal, Till, et al.
Publicado: (2024)
C2CServe: Leveraging NVLink-C2C for Elastic Serverless LLM Serving on MIG
por: Luo, Shutian, et al.
Publicado: (2026)
por: Luo, Shutian, et al.
Publicado: (2026)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
por: Görgen, Jakob, et al.
Publicado: (2024)
por: Görgen, Jakob, et al.
Publicado: (2024)
AdaptCache: KV Cache Native Storage Hierarchy for Low-Delay and High-Quality Language Model Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
Idiosyncrasies of Programmable Caching Engines
por: Peixoto, José, et al.
Publicado: (2026)
por: Peixoto, José, et al.
Publicado: (2026)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
por: Cai, Tianhao, et al.
Publicado: (2025)
por: Cai, Tianhao, et al.
Publicado: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
por: Tofigh, Mani, et al.
Publicado: (2025)
por: Tofigh, Mani, et al.
Publicado: (2025)
Age-Memory Trade-off in Read-Copy-Update
por: Ramani, Vishakha, et al.
Publicado: (2024)
por: Ramani, Vishakha, et al.
Publicado: (2024)
ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
por: Wang, Shihao, et al.
Publicado: (2026)
por: Wang, Shihao, et al.
Publicado: (2026)
From Good to Great: Improving Memory Tiering Performance Through Parameter Tuning
por: Kanellis, Konstantinos, et al.
Publicado: (2025)
por: Kanellis, Konstantinos, et al.
Publicado: (2025)
Assessing FIFO and Round Robin Scheduling:Effects on Data Pipeline Performance and Energy Usage
por: Choudhury, Malobika Roy, et al.
Publicado: (2024)
por: Choudhury, Malobika Roy, et al.
Publicado: (2024)
Phoenix -- A Novel Technique for Performance-Aware Orchestration of Thread and Page Table Placement in NUMA Systems
por: Siavashi, Mohammad, et al.
Publicado: (2025)
por: Siavashi, Mohammad, et al.
Publicado: (2025)
Valve: Production Online-Offline Inference Colocation with Jointly-Bounded Preemption Latency and Rate
por: Liu, Fangyue, et al.
Publicado: (2026)
por: Liu, Fangyue, et al.
Publicado: (2026)
BLITZSCALE: Fast and Live Large Model Autoscaling with O(1) Host Caching
por: Zhang, Dingyan, et al.
Publicado: (2024)
por: Zhang, Dingyan, et al.
Publicado: (2024)
The First Principle of Big Memory Systems
por: Hua, Yu
Publicado: (2023)
por: Hua, Yu
Publicado: (2023)
EVICPRESS: Joint KV-Cache Compression and Eviction for Efficient LLM Serving
por: Feng, Shaoting, et al.
Publicado: (2025)
por: Feng, Shaoting, et al.
Publicado: (2025)
DPC: A Distributed Page Cache over CXL
por: Bergman, Shai, et al.
Publicado: (2026)
por: Bergman, Shai, et al.
Publicado: (2026)
RUISA Operational Ecosystem Architecture
por: AL Mohtar, Mouayad
Publicado: (2026)
por: AL Mohtar, Mouayad
Publicado: (2026)
LLM as a System Service on Mobile Devices
por: Yin, Wangsong, et al.
Publicado: (2024)
por: Yin, Wangsong, et al.
Publicado: (2024)
I/O Transit Caching for PMem-based Block Device
por: Xu, Qing, et al.
Publicado: (2024)
por: Xu, Qing, et al.
Publicado: (2024)
Qurator: Scheduling Hybrid Quantum-Classical Workflows Across Heterogeneous Cloud Providers
por: Pehlivanoglu, Sinan, et al.
Publicado: (2026)
por: Pehlivanoglu, Sinan, et al.
Publicado: (2026)
Puzzle: Scheduling Multiple Deep Learning Models on Mobile Device with Heterogeneous Processors
por: Kang, Duseok, et al.
Publicado: (2025)
por: Kang, Duseok, et al.
Publicado: (2025)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
por: Moura, Diego, et al.
Publicado: (2022)
por: Moura, Diego, et al.
Publicado: (2022)
LearnedFTL: A Learning-Based Page-Level FTL for Reducing Double Reads in Flash-Based SSDs
por: Wang, Shengzhe, et al.
Publicado: (2023)
por: Wang, Shengzhe, et al.
Publicado: (2023)
Ejemplares similares
-
ConZone+: Practical Zoned Flash Storage Emulation for Consumer Devices
por: Yu, Dingcui, et al.
Publicado: (2025) -
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
por: Qiu, Shi, et al.
Publicado: (2026) -
Cache is King: Smart Page Eviction with eBPF
por: Zussman, Tal, et al.
Publicado: (2025) -
Guidelines for Building Indexes on Partially Cache-Coherent CXL Shared Memory
por: Wu, Fangnuo, et al.
Publicado: (2025) -
Victima: Drastically Increasing Address Translation Reach by Leveraging Underutilized Cache Resources
por: Kanellopoulos, Konstantinos, et al.
Publicado: (2023)