Managed-Retention Memory: A New Class of Memory for the AI Era
Fuente:
arXiv
Guardado en:
| Autores principales: | Legtchenko, Sergey, Stefanovici, Ioan, Black, Richard, Rowstron, Antony, Liu, Junyi, Costa, Paolo, Canakci, Burcu, Narayanan, Dushyanth, Wu, Xingbo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
por: Canakci, Burcu, et al.
Publicado: (2025)
por: Canakci, Burcu, et al.
Publicado: (2025)
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
por: Call, Aaron, et al.
Publicado: (2025)
por: Call, Aaron, et al.
Publicado: (2025)
CLAASIC: a Cortex-Inspired Hardware Accelerator
por: Puente, Valentin, et al.
Publicado: (2016)
por: Puente, Valentin, et al.
Publicado: (2016)
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
por: Zhang, Xu, et al.
Publicado: (2024)
por: Zhang, Xu, et al.
Publicado: (2024)
Harnessing the Full Potential of RRAMs through Scalable and Distributed In-Memory Computing with Integrated Error Correction
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
por: Vo, Huynh Q. N., et al.
Publicado: (2025)
Wattlytics: A Web Platform for Co-Optimizing Performance, Energy, and TCO in HPC Clusters
por: Afzal, Ayesha, et al.
Publicado: (2026)
por: Afzal, Ayesha, et al.
Publicado: (2026)
COMPASS: A Compiler Framework for Resource-Constrained Crossbar-Array Based In-Memory Deep Learning Accelerators
por: Park, Jihoon, et al.
Publicado: (2025)
por: Park, Jihoon, et al.
Publicado: (2025)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
por: Shi, Man, et al.
Publicado: (2024)
por: Shi, Man, et al.
Publicado: (2024)
Memory-Centric Computing: Solving Computing's Memory Problem
por: Mutlu, Onur, et al.
Publicado: (2025)
por: Mutlu, Onur, et al.
Publicado: (2025)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
por: Volos, Haris, et al.
Publicado: (2025)
por: Volos, Haris, et al.
Publicado: (2025)
TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms
por: Hou, Yuewen, et al.
Publicado: (2025)
por: Hou, Yuewen, et al.
Publicado: (2025)
Architecting Distributed Quantum Computers: Design Insights from Resource Estimation
por: Filippov, Dmitry, et al.
Publicado: (2025)
por: Filippov, Dmitry, et al.
Publicado: (2025)
ForgetMeNot: Understanding and Modeling the Impact of Forever Chemicals Toward Sustainable Large-Scale Computing
por: Roy, Rohan Basu, et al.
Publicado: (2025)
por: Roy, Rohan Basu, et al.
Publicado: (2025)
Carbon Connect: An Ecosystem for Sustainable Computing
por: Lee, Benjamin C., et al.
Publicado: (2024)
por: Lee, Benjamin C., et al.
Publicado: (2024)
Reference Architecture of a Quantum-Centric Supercomputer
por: Seelam, Seetharami, et al.
Publicado: (2026)
por: Seelam, Seetharami, et al.
Publicado: (2026)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
por: Ortega, Cristobal, et al.
Publicado: (2024)
por: Ortega, Cristobal, et al.
Publicado: (2024)
A Modern Primer on Processing in Memory
por: Mutlu, Onur, et al.
Publicado: (2020)
por: Mutlu, Onur, et al.
Publicado: (2020)
PIUMA: Programmable Integrated Unified Memory Architecture
por: Aananthakrishnan, Sriram, et al.
Publicado: (2020)
por: Aananthakrishnan, Sriram, et al.
Publicado: (2020)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
por: Asquini, Lorenzo, et al.
Publicado: (2025)
por: Asquini, Lorenzo, et al.
Publicado: (2025)
Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
por: Ibrahim, Mohamed Assem, et al.
Publicado: (2024)
Efficient Architecture for RISC-V Vector Memory Access
por: Guan, Hongyi, et al.
Publicado: (2025)
por: Guan, Hongyi, et al.
Publicado: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
por: Mutlu, Onur, et al.
Publicado: (2024)
por: Mutlu, Onur, et al.
Publicado: (2024)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
por: Li, Jiamin, et al.
Publicado: (2025)
por: Li, Jiamin, et al.
Publicado: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
por: Psistakis, Antonis
Publicado: (2025)
por: Psistakis, Antonis
Publicado: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
por: Ma, Ruiyang, et al.
Publicado: (2026)
por: Ma, Ruiyang, et al.
Publicado: (2026)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
por: Oliveira, Geraldo F.
Publicado: (2025)
por: Oliveira, Geraldo F.
Publicado: (2025)
TeraPool: A Physical Design Aware, 1024 RISC-V Cores Shared-L1-Memory Scaled-up Cluster Design with High Bandwidth Main Memory Link
por: Zhang, Yichao, et al.
Publicado: (2026)
por: Zhang, Yichao, et al.
Publicado: (2026)
Towards Memory Specialization: A Case for Long-Term and Short-Term RAM
por: Li, Peijing, et al.
Publicado: (2025)
por: Li, Peijing, et al.
Publicado: (2025)
BlockAMC: Scalable In-Memory Analog Matrix Computing for Solving Linear Systems
por: Pan, Lunshuai, et al.
Publicado: (2024)
por: Pan, Lunshuai, et al.
Publicado: (2024)
Survey of Disaggregated Memory: Cross-layer Technique Insights for Next-Generation Datacenters
por: Wang, Jing, et al.
Publicado: (2025)
por: Wang, Jing, et al.
Publicado: (2025)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
por: Frouzakis, Manos, et al.
Publicado: (2025)
por: Frouzakis, Manos, et al.
Publicado: (2025)
A Programming Model for Disaggregated Memory over CXL
por: Assa, Gal, et al.
Publicado: (2024)
por: Assa, Gal, et al.
Publicado: (2024)
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models
por: Tithi, Jesmin Jahan, et al.
Publicado: (2025)
por: Tithi, Jesmin Jahan, et al.
Publicado: (2025)
Efficient Optimization Accelerator Framework for Multistate Ising Problems
por: Garg, Chirag, et al.
Publicado: (2025)
por: Garg, Chirag, et al.
Publicado: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
por: Liu, Lian, et al.
Publicado: (2026)
por: Liu, Lian, et al.
Publicado: (2026)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
por: Yu, Yanpeng, et al.
Publicado: (2025)
por: Yu, Yanpeng, et al.
Publicado: (2025)
SpeedMalloc: Improving Multi-threaded Applications via a Lightweight Core for Memory Allocation
por: Li, Ruihao, et al.
Publicado: (2025)
por: Li, Ruihao, et al.
Publicado: (2025)
CCCL: Node-Spanning GPU Collectives with CXL Memory Pooling
por: Xu, Dong, et al.
Publicado: (2026)
por: Xu, Dong, et al.
Publicado: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
por: Chen, Yanru, et al.
Publicado: (2025)
por: Chen, Yanru, et al.
Publicado: (2025)
Ejemplares similares
-
Good things come in small packages: Should we build AI clusters with Lite-GPUs?
por: Canakci, Burcu, et al.
Publicado: (2025) -
From GPUs to RRAMs: Distributed In-Memory Primal-Dual Hybrid Gradient Method for Solving Large-Scale Linear Optimization Problem
por: Vo, Huynh Q. N., et al.
Publicado: (2025) -
Open Challenges for a Production-ready Cloud Environment on top of RISC-V hardware
por: Call, Aaron, et al.
Publicado: (2025) -
CLAASIC: a Cortex-Inspired Hardware Accelerator
por: Puente, Valentin, et al.
Publicado: (2016) -
DFabric: Scaling Out Data Parallel Applications with CXL-Ethernet Hybrid Interconnects
por: Zhang, Xu, et al.
Publicado: (2024)