Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
Fuente:
arXiv
Salvato in:
| Autori principali: | Hong, Jeongmin, Cho, Sungjun, Park, Geonwoo, Yang, Wonhyuk, Gong, Young-Ho, Kim, Gwangsun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
Exploring DRAM Cache Prefetching for Pooled Memory
di: Tirumalasetty, Chandrahas, et al.
Pubblicazione: (2024)
di: Tirumalasetty, Chandrahas, et al.
Pubblicazione: (2024)
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
TDRAM: Tag-enhanced DRAM for Efficient Caching
di: Babaie, Maryam, et al.
Pubblicazione: (2024)
di: Babaie, Maryam, et al.
Pubblicazione: (2024)
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
di: Nam, Hwayong, et al.
Pubblicazione: (2024)
di: Nam, Hwayong, et al.
Pubblicazione: (2024)
Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-aware Cache Compression
di: Cheng, Feng, et al.
Pubblicazione: (2025)
di: Cheng, Feng, et al.
Pubblicazione: (2025)
LOCALUT: Harnessing Capacity-Computation Tradeoffs for LUT-Based Inference in DRAM-PIM
di: Hong, Junguk, et al.
Pubblicazione: (2026)
di: Hong, Junguk, et al.
Pubblicazione: (2026)
Read Disturbance in High Bandwidth Memory: A Detailed Experimental Study on HBM2 DRAM Chips
di: Olgun, Ataberk, et al.
Pubblicazione: (2023)
di: Olgun, Ataberk, et al.
Pubblicazione: (2023)
AERO: Adaptive Erase Operation for Improving Lifetime and Performance of Modern NAND Flash-Based SSDs
di: Cho, Sungjun, et al.
Pubblicazione: (2024)
di: Cho, Sungjun, et al.
Pubblicazione: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
di: Heo, Guseul, et al.
Pubblicazione: (2024)
di: Heo, Guseul, et al.
Pubblicazione: (2024)
RED: Energy Optimization Framework for eDRAM-based PIM with Reconfigurable Voltage Swing and Retention-aware Scheduling
di: Kim, Jae-Young, et al.
Pubblicazione: (2025)
di: Kim, Jae-Young, et al.
Pubblicazione: (2025)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
di: Kim, Junsoo, et al.
Pubblicazione: (2025)
di: Kim, Junsoo, et al.
Pubblicazione: (2025)
Shifting in-DRAM
di: Tegge, William C., et al.
Pubblicazione: (2026)
di: Tegge, William C., et al.
Pubblicazione: (2026)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
di: Lu, Tsung-Han, et al.
Pubblicazione: (2026)
di: Lu, Tsung-Han, et al.
Pubblicazione: (2026)
Securing DRAM at Scale: ARFM-Driven Row Hammer Defense with Unveiling the Threat of Short tRC Patterns
di: Joo, Nogeun, et al.
Pubblicazione: (2025)
di: Joo, Nogeun, et al.
Pubblicazione: (2025)
Sudoku: Decomposing DRAM Address Mapping into Component Functions
di: Wi, Minbok, et al.
Pubblicazione: (2025)
di: Wi, Minbok, et al.
Pubblicazione: (2025)
Pushing the Memory Bandwidth Wall with CXL-enabled Idle I/O Bandwidth Harvesting
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
Per-Bank Bandwidth Regulation of Shared Last-Level Cache for Real-Time Systems
di: Sullivan, Connor, et al.
Pubblicazione: (2024)
di: Sullivan, Connor, et al.
Pubblicazione: (2024)
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture
di: Olgun, Ataberk, et al.
Pubblicazione: (2022)
di: Olgun, Ataberk, et al.
Pubblicazione: (2022)
SAL-PIM: A Subarray-level Processing-in-Memory Architecture with LUT-based Linear Interpolation for Transformer-based Text Generation
di: Han, Wontak, et al.
Pubblicazione: (2024)
di: Han, Wontak, et al.
Pubblicazione: (2024)
EasyDRAM: An FPGA-based Infrastructure for Fast and Accurate End-to-End Evaluation of Emerging DRAM Techniques
di: Canpolat, Oğuzhan, et al.
Pubblicazione: (2025)
di: Canpolat, Oğuzhan, et al.
Pubblicazione: (2025)
System-Technology Co-Optimization of Bitline Routing and Bonding Pathways in Monolithic 3D DRAM Architectures
di: Lee, Kiseok, et al.
Pubblicazione: (2026)
di: Lee, Kiseok, et al.
Pubblicazione: (2026)
Area-Efficient In-Memory Computing for Mixture-of-Experts via Multiplexing and Caching
di: Gao, Hanyuan, et al.
Pubblicazione: (2026)
di: Gao, Hanyuan, et al.
Pubblicazione: (2026)
Scaling Routers with In-Package Optics and High-Bandwidth Memories
di: Keslassy, Isaac, et al.
Pubblicazione: (2026)
di: Keslassy, Isaac, et al.
Pubblicazione: (2026)
Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
di: Xia, Tianhua, et al.
Pubblicazione: (2025)
Block-SSD: A New Block-Based Blocking SSD Architecture
di: Wong, Ryan, et al.
Pubblicazione: (2024)
di: Wong, Ryan, et al.
Pubblicazione: (2024)
Generalized Ping-Pong: Off-Chip Memory Bandwidth Centric Pipelining Strategy for Processing-In-Memory Accelerators
di: Wang, Ruibao, et al.
Pubblicazione: (2024)
di: Wang, Ruibao, et al.
Pubblicazione: (2024)
GEM3D CIM General Purpose Matrix Computation Using 3D Integrated SRAM eDRAM Hybrid Compute In Memory on Memory Architecture
di: Chakraborty, Subhradip, et al.
Pubblicazione: (2026)
di: Chakraborty, Subhradip, et al.
Pubblicazione: (2026)
Per-Bank Memory Bandwidth Regulation for Predictable and Performant Real-Time System
di: Sullivan, Connor Rudy, et al.
Pubblicazione: (2026)
di: Sullivan, Connor Rudy, et al.
Pubblicazione: (2026)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
di: Xie, Rui, et al.
Pubblicazione: (2025)
di: Xie, Rui, et al.
Pubblicazione: (2025)
Improving the Representativeness of Simulation Intervals for the Cache Memory System
di: Bueno, Nicolas, et al.
Pubblicazione: (2024)
di: Bueno, Nicolas, et al.
Pubblicazione: (2024)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
di: Kim, Minsu, et al.
Pubblicazione: (2025)
di: Kim, Minsu, et al.
Pubblicazione: (2025)
Rethinking the Producer-Consumer Relationship in Modern DRAM-Based Systems
di: Patel, Minesh, et al.
Pubblicazione: (2024)
di: Patel, Minesh, et al.
Pubblicazione: (2024)
DRAM-Profiler: An Experimental DRAM RowHammer Vulnerability Profiling Mechanism
di: Zhou, Ranyang, et al.
Pubblicazione: (2024)
di: Zhou, Ranyang, et al.
Pubblicazione: (2024)
CMD: A Cache-assisted GPU Memory Deduplication Architecture
di: Zhao, Wei, et al.
Pubblicazione: (2024)
di: Zhao, Wei, et al.
Pubblicazione: (2024)
Virgo: Cluster-level Matrix Unit Integration in GPUs for Scalability and Energy Efficiency
di: Kim, Hansung, et al.
Pubblicazione: (2024)
di: Kim, Hansung, et al.
Pubblicazione: (2024)
Pushing up to the Limit of Memory Bandwidth and Capacity Utilization for Efficient LLM Decoding on Embedded FPGA
di: Li, Jindong, et al.
Pubblicazione: (2025)
di: Li, Jindong, et al.
Pubblicazione: (2025)
Low-Power Encoding for PAM-3 DRAM Bus
di: Nam, Jonghyeon, et al.
Pubblicazione: (2024)
di: Nam, Jonghyeon, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024) -
Exploring DRAM Cache Prefetching for Pooled Memory
di: Tirumalasetty, Chandrahas, et al.
Pubblicazione: (2024) -
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024) -
TDRAM: Tag-enhanced DRAM for Efficient Caching
di: Babaie, Maryam, et al.
Pubblicazione: (2024) -
DRAMScope: Uncovering DRAM Microarchitecture and Characteristics by Issuing Memory Commands
di: Nam, Hwayong, et al.
Pubblicazione: (2024)