Managed-Retention Memory: A New Class of Memory for the AI Era
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Legtchenko, Sergey, Stefanovici, Ioan, Black, Richard, Rowstron, Antony, Liu, Junyi, Costa, Paolo, Canakci, Burcu, Narayanan, Dushyanth, Wu, Xingbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
von: Canakci, Burcu, et al.
Veröffentlicht: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
von: Call, Aaron, et al.
Veröffentlicht: (2025)
von: Call, Aaron, et al.
Veröffentlicht: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
von: Puente, Valentin, et al.
Veröffentlicht: (2016)
von: Puente, Valentin, et al.
Veröffentlicht: (2016)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
von: Zhang, Xu, et al.
Veröffentlicht: (2024)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
von: Afzal, Ayesha, et al.
Veröffentlicht: (2026)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
von: Park, Jihoon, et al.
Veröffentlicht: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
von: Shi, Man, et al.
Veröffentlicht: (2024)
von: Shi, Man, et al.
Veröffentlicht: (2024)
Memory-Centric Computing: Solving Computing's Memory Problem
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
von: Volos, Haris, et al.
Veröffentlicht: (2025)
von: Volos, Haris, et al.
Veröffentlicht: (2025)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
von: Hou, Yuewen, et al.
Veröffentlicht: (2025)
von: Hou, Yuewen, et al.
Veröffentlicht: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
von: Filippov, Dmitry, et al.
Veröffentlicht: (2025)
von: Filippov, Dmitry, et al.
Veröffentlicht: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
von: Roy, Rohan Basu, et al.
Veröffentlicht: (2025)
von: Roy, Rohan Basu, et al.
Veröffentlicht: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
von: Lee, Benjamin C., et al.
Veröffentlicht: (2024)
von: Lee, Benjamin C., et al.
Veröffentlicht: (2024)
Reference Architecture of a Quantum-Centric Supercomputer
von: Seelam, Seetharami, et al.
Veröffentlicht: (2026)
von: Seelam, Seetharami, et al.
Veröffentlicht: (2026)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
von: Ortega, Cristobal, et al.
Veröffentlicht: (2024)
von: Ortega, Cristobal, et al.
Veröffentlicht: (2024)
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
PIUMA: Programmable Integrated Unified Memory Architecture
von: Aananthakrishnan, Sriram, et al.
Veröffentlicht: (2020)
von: Aananthakrishnan, Sriram, et al.
Veröffentlicht: (2020)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
von: Ibrahim, Mohamed Assem, et al.
Veröffentlicht: (2024)
Efficient Architecture for RISC-V Vector Memory Access
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
von: Guan, Hongyi, et al.
Veröffentlicht: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
von: Li, Jiamin, et al.
Veröffentlicht: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
von: Psistakis, Antonis
Veröffentlicht: (2025)
von: Psistakis, Antonis
Veröffentlicht: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
von: Ma, Ruiyang, et al.
Veröffentlicht: (2026)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
von: Oliveira, Geraldo F.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F.
Veröffentlicht: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
von: Zhang, Yichao, et al.
Veröffentlicht: (2026)
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
von: Li, Peijing, et al.
Veröffentlicht: (2025)
von: Li, Peijing, et al.
Veröffentlicht: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
von: Pan, Lunshuai, et al.
Veröffentlicht: (2024)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
von: Wang, Jing, et al.
Veröffentlicht: (2025)
von: Wang, Jing, et al.
Veröffentlicht: (2025)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
A Programming Model for Disaggregated Memory over CXL
von: Assa, Gal, et al.
Veröffentlicht: (2024)
von: Assa, Gal, et al.
Veröffentlicht: (2024)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2025)
von: Tithi, Jesmin Jahan, et al.
Veröffentlicht: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
von: Garg, Chirag, et al.
Veröffentlicht: (2025)
von: Garg, Chirag, et al.
Veröffentlicht: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
von: Liu, Lian, et al.
Veröffentlicht: (2026)
von: Liu, Lian, et al.
Veröffentlicht: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
von: Yu, Yanpeng, et al.
Veröffentlicht: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
von: Li, Ruihao, et al.
Veröffentlicht: (2025)
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
von: Xu, Dong, et al.
Veröffentlicht: (2026)
von: Xu, Dong, et al.
Veröffentlicht: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
von: Canakci, Burcu, et al.
Veröffentlicht: (2025) -
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
von: Vo, Huynh Q. N., et al.
Veröffentlicht: (2025) -
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
von: Call, Aaron, et al.
Veröffentlicht: (2025) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
von: Puente, Valentin, et al.
Veröffentlicht: (2016) -
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
von: Zhang, Xu, et al.
Veröffentlicht: (2024)