CacheMind: From Miss Rates to Why -- Natural-Language, Trace-Grounded Reasoning for Cache Replacement
Fuente:
arXiv
Guardado en:
| Autores principales: | Mhapsekar, Kaushal, Ghanbari, Azam, Aslrousta, Bita, Mirbagher-Ajorpaz, Samira |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX
por: Dizani, Farshad, et al.
Publicado: (2025)
por: Dizani, Farshad, et al.
Publicado: (2025)
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
por: Choi, Yuseon, et al.
Publicado: (2025)
por: Choi, Yuseon, et al.
Publicado: (2025)
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
por: You, Dean, et al.
Publicado: (2025)
por: You, Dean, et al.
Publicado: (2025)
Comparative Characterization of KV Cache Management Strategies for LLM Inference
por: Mamo, Oteo, et al.
Publicado: (2026)
por: Mamo, Oteo, et al.
Publicado: (2026)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
por: Xia, Tianhua, et al.
Publicado: (2025)
por: Xia, Tianhua, et al.
Publicado: (2025)
BackCache: Mitigating Contention-Based Cache Timing Attacks by Hiding Cache Line Evictions
por: Wang, Quancheng, et al.
Publicado: (2023)
por: Wang, Quancheng, et al.
Publicado: (2023)
Adaptive Cache Pollution Control for Large Language Model Inference Workloads Using Temporal CNN-Based Prediction and Priority-Aware Replacement
por: Liu, Songze, et al.
Publicado: (2025)
por: Liu, Songze, et al.
Publicado: (2025)
Exploring DRAM Cache Prefetching for Pooled Memory
por: Tirumalasetty, Chandrahas, et al.
Publicado: (2024)
por: Tirumalasetty, Chandrahas, et al.
Publicado: (2024)
Potential and Limitation of High-Frequency Cores and Caches
por: Pai, Kunal, et al.
Publicado: (2024)
por: Pai, Kunal, et al.
Publicado: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
por: Babaie, Maryam, et al.
Publicado: (2024)
por: Babaie, Maryam, et al.
Publicado: (2024)
The Avatar Cache: Enabling On-Demand Security with Morphable Cache Architecture
por: Bhatla, Anubhav, et al.
Publicado: (2026)
por: Bhatla, Anubhav, et al.
Publicado: (2026)
Random and Safe Cache Architecture to Defeat Cache Timing Attacks
por: Hu, Guangyuan, et al.
Publicado: (2023)
por: Hu, Guangyuan, et al.
Publicado: (2023)
Systematic Evaluation of Randomized Cache Designs against Cache Occupancy
por: Chakraborty, Anirban, et al.
Publicado: (2023)
por: Chakraborty, Anirban, et al.
Publicado: (2023)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
por: Fang, Yunhua, et al.
Publicado: (2025)
por: Fang, Yunhua, et al.
Publicado: (2025)
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
por: Skliar, Andrii, et al.
Publicado: (2024)
por: Skliar, Andrii, et al.
Publicado: (2024)
Fletch: File-System Metadata Caching in Programmable Switches
por: Liu, Qingxiu, et al.
Publicado: (2025)
por: Liu, Qingxiu, et al.
Publicado: (2025)
Improving the Representativeness of Simulation Intervals for the Cache Memory System
por: Bueno, Nicolas, et al.
Publicado: (2024)
por: Bueno, Nicolas, et al.
Publicado: (2024)
VeriCache: Turning Lossy KV Cache into Lossless LLM Inference
por: Yao, Jiayi, et al.
Publicado: (2026)
por: Yao, Jiayi, et al.
Publicado: (2026)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
por: Kim, Dowon, et al.
Publicado: (2025)
por: Kim, Dowon, et al.
Publicado: (2025)
Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
por: Khadem, Alireza, et al.
Publicado: (2025)
por: Khadem, Alireza, et al.
Publicado: (2025)
Pickle Prefetcher: Programmable and Scalable Last-Level Cache Prefetcher
por: Nguyen, Hoa, et al.
Publicado: (2025)
por: Nguyen, Hoa, et al.
Publicado: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
por: Zhao, Wei, et al.
Publicado: (2024)
por: Zhao, Wei, et al.
Publicado: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
por: Hong, Jeongmin, et al.
Publicado: (2024)
por: Hong, Jeongmin, et al.
Publicado: (2024)
RollingCache: Using Runtime Behavior to Defend Against Cache Side Channel Attacks
por: Ojha, Divya, et al.
Publicado: (2024)
por: Ojha, Divya, et al.
Publicado: (2024)
CaMDN: Enhancing Cache Efficiency for Multi-tenant DNNs on Integrated NPUs
por: Cai, Tianhao, et al.
Publicado: (2025)
por: Cai, Tianhao, et al.
Publicado: (2025)
ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
por: Xing, Changwen, et al.
Publicado: (2025)
por: Xing, Changwen, et al.
Publicado: (2025)
3D MPSoC with On-Chip Cache Support -- Design and Exploitation
por: Cataldo, Rodrigo, et al.
Publicado: (2025)
por: Cataldo, Rodrigo, et al.
Publicado: (2025)
An Analytical and Empirical Investigation of Tag Partitioning for Energy-Efficient Reliable Cache
por: Cheshmikhani, Elham, et al.
Publicado: (2025)
por: Cheshmikhani, Elham, et al.
Publicado: (2025)
On the Impact of ISA Extension on Energy Consumption of I-Cache in Extensible Processors
por: Behboudi, Noushin, et al.
Publicado: (2024)
por: Behboudi, Noushin, et al.
Publicado: (2024)
Titanus: Enabling KV Cache Pruning and Quantization On-the-Fly for LLM Acceleration
por: Chen, Peilin, et al.
Publicado: (2025)
por: Chen, Peilin, et al.
Publicado: (2025)
ARCANE: Adaptive RISC-V Cache Architecture for Near-memory Extensions
por: Petrolo, Vincenzo, et al.
Publicado: (2025)
por: Petrolo, Vincenzo, et al.
Publicado: (2025)
Enhancing Reliability of STT-MRAM Caches by Eliminating Read Disturbance Accumulation
por: Cheshmikhani, Elham, et al.
Publicado: (2026)
por: Cheshmikhani, Elham, et al.
Publicado: (2026)
The Open-Source BlackParrot-BedRock Cache Coherence System
por: Wyse, Mark Unruh
Publicado: (2025)
por: Wyse, Mark Unruh
Publicado: (2025)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
por: Gao, Hanyuan, et al.
Publicado: (2026)
por: Gao, Hanyuan, et al.
Publicado: (2026)
Learning Cache Coherence Traffic for NoC Routing Design
por: Xiong, Guochu, et al.
Publicado: (2025)
por: Xiong, Guochu, et al.
Publicado: (2025)
ROBIN: Incremental Oblique Interleaved ECC for Reliability Improvement in STT-MRAM Caches
por: Cheshmikhani, Elham, et al.
Publicado: (2026)
por: Cheshmikhani, Elham, et al.
Publicado: (2026)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
por: Nguyen, Hoa, et al.
Publicado: (2025)
por: Nguyen, Hoa, et al.
Publicado: (2025)
Enhancing Instruction Prefetching via Cache and TLB Management
por: Jamet, Alexandre Valentin, et al.
Publicado: (2026)
por: Jamet, Alexandre Valentin, et al.
Publicado: (2026)
The Bicameral Cache: a split cache for vector architectures
por: Rebolledo, Susana, et al.
Publicado: (2024)
por: Rebolledo, Susana, et al.
Publicado: (2024)
PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
por: Li, Luyi, et al.
Publicado: (2023)
por: Li, Luyi, et al.
Publicado: (2023)
Ejemplares similares
-
THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX
por: Dizani, Farshad, et al.
Publicado: (2025) -
SliceMoE: Bit-Sliced Expert Caching under Miss-Rate Constraints for Efficient MoE Inference
por: Choi, Yuseon, et al.
Publicado: (2025) -
MERE: Hardware-Software Co-Design for Masking Cache Miss Latency in Embedded Processors
por: You, Dean, et al.
Publicado: (2025) -
Comparative Characterization of KV Cache Management Strategies for LLM Inference
por: Mamo, Oteo, et al.
Publicado: (2026) -
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
por: Xia, Tianhua, et al.
Publicado: (2025)