Prefetching in Deep Memory Hierarchies with NVRAM as Main Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Lurbe, Manel, Avargues, Miguel, Petit, Salvador, Gomez, Maria E., Yang, Rui, Wang, Guanhao, Sahuquillo, Julio |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
by: Li, Zhilin, et al.
Published: (2025)
by: Li, Zhilin, et al.
Published: (2025)
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
by: Navarro, Marta, et al.
Published: (2025)
by: Navarro, Marta, et al.
Published: (2025)
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025)
by: Huang, Yutong, et al.
Published: (2025)
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026)
by: Wang, Wenfeng, et al.
Published: (2026)
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
by: Guo, Cong, et al.
Published: (2024)
by: Guo, Cong, et al.
Published: (2024)
DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance
by: Zhang, Yuning, et al.
Published: (2025)
by: Zhang, Yuning, et al.
Published: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026)
by: Liu, Lian, et al.
Published: (2026)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
SP-MoE: Speculative Decoding and Prefetching for Accelerating MoE-based Model Inference
by: Chen, Liangkun, et al.
Published: (2025)
by: Chen, Liangkun, et al.
Published: (2025)
On the Bit Complexity of Iterated Memory
by: Toyos-Marfurt, Guillermo, et al.
Published: (2024)
by: Toyos-Marfurt, Guillermo, et al.
Published: (2024)
eLLM: Elastic Memory Management Framework for Efficient LLM Serving
by: Xu, Jiale, et al.
Published: (2025)
by: Xu, Jiale, et al.
Published: (2025)
FlashMem: Supporting Modern DNN Workloads on Mobile with GPU Memory Hierarchy Optimizations
by: Shu, Zhihao, et al.
Published: (2026)
by: Shu, Zhihao, et al.
Published: (2026)
PROBE: Co-Balancing Computation and Communication in MoE Inference via Real-Time Predictive Prefetching
by: Zhu, Qianchao, et al.
Published: (2026)
by: Zhu, Qianchao, et al.
Published: (2026)
Memory Offloading for Large Language Model Inference with Latency SLO Guarantees
by: Ma, Chenxiang, et al.
Published: (2025)
by: Ma, Chenxiang, et al.
Published: (2025)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
by: Ghiasi, Nika Mansouri, et al.
Published: (2022)
LOCO: Rethinking Objects for Network Memory
by: Hodgkins, George, et al.
Published: (2025)
by: Hodgkins, George, et al.
Published: (2025)
What Cannot Be Implemented on Weak Memory?
by: Castañeda, Armando, et al.
Published: (2024)
by: Castañeda, Armando, et al.
Published: (2024)
Self-Evolving Distributed Memory Architecture for Scalable AI Systems
by: Li, Zixuan, et al.
Published: (2026)
by: Li, Zixuan, et al.
Published: (2026)
BladeDISC++: Memory Optimizations Based On Symbolic Shape
by: Yuan, Xiulong, et al.
Published: (2024)
by: Yuan, Xiulong, et al.
Published: (2024)
Deterministic Self-Stabilising Leader Election for Programmable Matter with Constant Memory
by: Chalopin, Jérémie, et al.
Published: (2024)
by: Chalopin, Jérémie, et al.
Published: (2024)
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
by: Zhang, Chen, et al.
Published: (2025)
by: Zhang, Chen, et al.
Published: (2025)
FlexKV: Flexible Index Offloading for Memory-Disaggregated Key-Value Store
by: Hu, Zhisheng, et al.
Published: (2025)
by: Hu, Zhisheng, et al.
Published: (2025)
Generalized Data Placement Strategies for Racetrack Memories
by: Khan, Asif Ali, et al.
Published: (2019)
by: Khan, Asif Ali, et al.
Published: (2019)
IM-PIR: In-Memory Private Information Retrieval
by: Mwaisela, Mpoki, et al.
Published: (2025)
by: Mwaisela, Mpoki, et al.
Published: (2025)
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
by: Yang, Shuo, et al.
Published: (2026)
by: Yang, Shuo, et al.
Published: (2026)
Breaking the Memory Wall: A Study of I/O Patterns and GPU Memory Utilization for Hybrid CPU-GPU Offloaded Optimizers
by: Maurya, Avinash, et al.
Published: (2024)
by: Maurya, Avinash, et al.
Published: (2024)
Ponder: Online Prediction of Task Memory Requirements for Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2024)
by: Lehmann, Fabian, et al.
Published: (2024)
FusionRCG: Orchestrating Recursive Computation Graphs across GPU Memory Hierarchies
by: Zhang, Yihong, et al.
Published: (2026)
by: Zhang, Yihong, et al.
Published: (2026)
Enhancing Memory Efficiency in Large Language Model Training Through Chronos-aware Pipeline Parallelism
by: Lin, Xinyuan, et al.
Published: (2025)
by: Lin, Xinyuan, et al.
Published: (2025)
Chameleon: Taming Dynamic Operator Sequences for Memory-Intensive LLM Training
by: Wang, Zibo, et al.
Published: (2025)
by: Wang, Zibo, et al.
Published: (2025)
PATSMA: Parameter Auto-tuning for Shared Memory Algorithms
by: Fernandes, Joao B., et al.
Published: (2024)
by: Fernandes, Joao B., et al.
Published: (2024)
DRackSim: Simulator for Rack-scale Memory Disaggregation
by: Puri, Amit, et al.
Published: (2023)
by: Puri, Amit, et al.
Published: (2023)
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
by: Bader, Jonathan, et al.
Published: (2024)
by: Bader, Jonathan, et al.
Published: (2024)
Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling
by: Bader, Jonathan, et al.
Published: (2026)
by: Bader, Jonathan, et al.
Published: (2026)
Gathering Teams of Bounded Memory Agents on a Line
by: Gao, Younan, et al.
Published: (2025)
by: Gao, Younan, et al.
Published: (2025)
Thread and Data Mapping in Software Transactional Memory: An Overview
by: Pasqualin, Douglas Pereira, et al.
Published: (2022)
by: Pasqualin, Douglas Pereira, et al.
Published: (2022)
Byzantine-Tolerant Consensus in GPU-Inspired Shared Memory
by: Georgiou, Chryssis, et al.
Published: (2025)
by: Georgiou, Chryssis, et al.
Published: (2025)
System-Level Performance Modeling of Photonic In-Memory Computing
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
by: Arockiaraj, Jebacyril, et al.
Published: (2026)
Similar Items
-
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
by: Li, Zhilin, et al.
Published: (2025) -
A New Family of Thread to Core Allocation Policies for an SMT ARM Processor
by: Navarro, Marta, et al.
Published: (2025) -
Learning Semantics, Not Addresses: Runtime Neural Prefetching for Far Memory
by: Huang, Yutong, et al.
Published: (2025) -
PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving
by: Wang, Wenfeng, et al.
Published: (2026) -
ML-based Adaptive Prefetching and Data Placement for US HEP Systems
by: Karanam, Venkat Sai Suman Lamba, et al.
Published: (2025)