Five-Minute Rule 40 Years Later: A First-Principles Revisit for Modern Memory Hierarchy
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Tong, Mailthody, Vikram Sharma, Sun, Fei, Ma, Linsen, Newburn, Chris J., Zhang, Teresa, Liu, Yang, Li, Jiangpeng, Zhong, Hao, Hwu, Wen-Mei |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
par: Park, Jeongmin Brian, et autres
Publié: (2023)
par: Park, Jeongmin Brian, et autres
Publié: (2023)
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Multiport Support for Vortex OpenGPU Memory Hierarchy
par: Shin, Injae, et autres
Publié: (2025)
par: Shin, Injae, et autres
Publié: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025)
par: Fang, Yunhua, et autres
Publié: (2025)
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Making Strong Error-Correcting Codes Work Effectively for HBM in AI Inference
par: Xie, Rui, et autres
Publié: (2025)
par: Xie, Rui, et autres
Publié: (2025)
Apparate: Evading Memory Hierarchy with GodSpeed Wireless-on-Chip
par: GS, Nitesh Narayana, et autres
Publié: (2024)
par: GS, Nitesh Narayana, et autres
Publié: (2024)
Theodosian: A Deep Dive into Memory-Hierarchy-Centric FHE Acceleration
par: Choi, Wonseok, et autres
Publié: (2025)
par: Choi, Wonseok, et autres
Publié: (2025)
A Configurable and Efficient Memory Hierarchy for Neural Network Hardware Accelerator
par: Bause, Oliver, et autres
Publié: (2024)
par: Bause, Oliver, et autres
Publié: (2024)
Memory Hierarchy Design for Caching Middleware in the Age of NVM
par: Ghandeharizadeh, Shahram, et autres
Publié: (2025)
par: Ghandeharizadeh, Shahram, et autres
Publié: (2025)
Revisiting Main Memory-Based Covert and Side Channel Attacks in the Context of Processing-in-Memory
par: Bostanci, F. Nisa, et autres
Publié: (2024)
par: Bostanci, F. Nisa, et autres
Publié: (2024)
täkōFormal: Enabling Robust Software for Programmable Memory Hierarchies (Extended Version)
par: Srinivasan, Pranav, et autres
Publié: (2026)
par: Srinivasan, Pranav, et autres
Publié: (2026)
Hardware-Software Co-Design for Accelerating Transformer Inference Leveraging Compute-in-Memory
par: Kim, Dong Eun, et autres
Publié: (2025)
par: Kim, Dong Eun, et autres
Publié: (2025)
A Modern Primer on Processing in Memory
par: Mutlu, Onur, et autres
Publié: (2020)
par: Mutlu, Onur, et autres
Publié: (2020)
Performance Characterizations and Usage Guidelines of Samsung CXL Memory Module Hybrid Prototype
par: Zeng, Jianping, et autres
Publié: (2025)
par: Zeng, Jianping, et autres
Publié: (2025)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
par: Zhao, Wei, et autres
Publié: (2024)
par: Zhao, Wei, et autres
Publié: (2024)
Descriptor-Based Object-Aware Memory Systems: A Comprehensive Review
par: Tong, Dong
Publié: (2025)
par: Tong, Dong
Publié: (2025)
HCiM: ADC-Less Hybrid Analog-Digital Compute in Memory Accelerator for Deep Learning Workloads
par: Negi, Shubham, et autres
Publié: (2024)
par: Negi, Shubham, et autres
Publié: (2024)
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access
par: Wang, Luming, et autres
Publié: (2024)
par: Wang, Luming, et autres
Publié: (2024)
Choreographer: A Full-System Framework for Fine-Grained Tasks in Cache Hierarchies
par: Nguyen, Hoa, et autres
Publié: (2025)
par: Nguyen, Hoa, et autres
Publié: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
par: Liu, Lian, et autres
Publié: (2026)
par: Liu, Lian, et autres
Publié: (2026)
SmartQuant: CXL-based AI Model Store in Support of Runtime Configurable Weight Quantization
par: Xie, Rui, et autres
Publié: (2024)
par: Xie, Rui, et autres
Publié: (2024)
Memory-Guided Unified Hardware Accelerator for Mixed-Precision Scientific Computing
par: Wang, Chuanzhen, et autres
Publié: (2026)
par: Wang, Chuanzhen, et autres
Publié: (2026)
Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation
par: Pinckney, Nathaniel, et autres
Publié: (2024)
par: Pinckney, Nathaniel, et autres
Publié: (2024)
A Review of SRAM-based Compute-in-Memory Circuits
par: Yoshioka, Kentaro, et autres
Publié: (2024)
par: Yoshioka, Kentaro, et autres
Publié: (2024)
AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
par: Cheng, Feng, et autres
Publié: (2025)
par: Cheng, Feng, et autres
Publié: (2025)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
par: Mao, Shunyu, et autres
Publié: (2024)
par: Mao, Shunyu, et autres
Publié: (2024)
Enabling Efficient Hardware Acceleration of Hybrid Vision Transformer (ViT) Networks at the Edge
par: Dumoulin, Joren, et autres
Publié: (2025)
par: Dumoulin, Joren, et autres
Publié: (2025)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
par: Wang, Jiapin, et autres
Publié: (2024)
par: Wang, Jiapin, et autres
Publié: (2024)
Control Flow Management in Modern GPUs
par: Shoushtary, Mojtaba Abaie, et autres
Publié: (2024)
par: Shoushtary, Mojtaba Abaie, et autres
Publié: (2024)
Analyzing Modern NVIDIA GPU cores
par: Huerta, Rodrigo, et autres
Publié: (2025)
par: Huerta, Rodrigo, et autres
Publié: (2025)
Mainframe-Style Channel Controllers for Modern Disaggregated Memory Systems
par: Liu, Zikai, et autres
Publié: (2025)
par: Liu, Zikai, et autres
Publié: (2025)
An Event-Driven Spiking Compute-In-Memory Macro based on SOT-MRAM
par: Yu, Deyang, et autres
Publié: (2025)
par: Yu, Deyang, et autres
Publié: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
par: Zhou, Zhe, et autres
Publié: (2024)
par: Zhou, Zhe, et autres
Publié: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
par: Cheng, Feng, et autres
Publié: (2025)
par: Cheng, Feng, et autres
Publié: (2025)
No One-Size-Fits-All: A Workload-Driven Characterization of Bit-Parallel vs. Bit-Serial Data Layouts for Processing-using-Memory
par: Zhang, Jingyao, et autres
Publié: (2025)
par: Zhang, Jingyao, et autres
Publié: (2025)
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
par: Yuan, Aojie, et autres
Publié: (2026)
par: Yuan, Aojie, et autres
Publié: (2026)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
par: Ghiasi, Nika Mansouri, et autres
Publié: (2022)
par: Ghiasi, Nika Mansouri, et autres
Publié: (2022)
The Case for Replication-Aware Memory-Error Protection in Disaggregated Memory
par: Volos, Haris
Publié: (2023)
par: Volos, Haris
Publié: (2023)
Documents similaires
-
Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage Accesses
par: Park, Jeongmin Brian, et autres
Publié: (2023) -
Reimagining Memory Access for LLM Inference: Compression-Aware Memory Controller Design
par: Xie, Rui, et autres
Publié: (2025) -
Multiport Support for Vortex OpenGPU Memory Hierarchy
par: Shin, Injae, et autres
Publié: (2025) -
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
par: Xie, Rui, et autres
Publié: (2025) -
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System
par: Fang, Yunhua, et autres
Publié: (2025)