PIM or CXL-PIM? Understanding Architectural Trade-offs Through Large-Scale Benchmarking

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Lee, I-Ting, Wang, Bao-Kai, Chen, Liang-Chi, Lim, Wen Sheng, Chang, Da-Wei, Chang, Yu-Ming, Ho, Chieng-Chung
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914164043153408
author Lee, I-Ting
Wang, Bao-Kai
Chen, Liang-Chi
Lim, Wen Sheng
Chang, Da-Wei
Chang, Yu-Ming
Ho, Chieng-Chung
author_facet Lee, I-Ting
Wang, Bao-Kai
Chen, Liang-Chi
Lim, Wen Sheng
Chang, Da-Wei
Chang, Yu-Ming
Ho, Chieng-Chung
contents Processing-in-memory (PIM) reduces data movement by executing near memory, but our large-scale characterization on real PIM hardware shows that end-to-end performance is often limited by disjoint host and device address spaces that force explicit staging transfers. In contrast, CXL-PIM provides a unified address space and cache-coherent access at the cost of higher access latency. These opposing interface models create workload-dependent tradeoffs that are not captured by small-scale studies. This work presents a side-by-side, large-scale comparison of PIM and CXL-PIM using measurements from real PIM hardware and trace-driven CXL modeling. We identify when unified-address access amortizes link latency enough to overcome transfer bottlenecks, and when tightly coupled PIM remains preferable. Our results reveal phase- and dataset-size regimes in which the relative ranking between the two architectures reverses, offering practical guidance for future near-memory system design.
format Preprint
id arxiv_https___arxiv_org_abs_2511_14400
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle PIM or CXL-PIM? Understanding Architectural Trade-offs Through Large-Scale Benchmarking
Lee, I-Ting
Wang, Bao-Kai
Chen, Liang-Chi
Lim, Wen Sheng
Chang, Da-Wei
Chang, Yu-Ming
Ho, Chieng-Chung
Emerging Technologies
Performance
Processing-in-memory (PIM) reduces data movement by executing near memory, but our large-scale characterization on real PIM hardware shows that end-to-end performance is often limited by disjoint host and device address spaces that force explicit staging transfers. In contrast, CXL-PIM provides a unified address space and cache-coherent access at the cost of higher access latency. These opposing interface models create workload-dependent tradeoffs that are not captured by small-scale studies. This work presents a side-by-side, large-scale comparison of PIM and CXL-PIM using measurements from real PIM hardware and trace-driven CXL modeling. We identify when unified-address access amortizes link latency enough to overcome transfer bottlenecks, and when tightly coupled PIM remains preferable. Our results reveal phase- and dataset-size regimes in which the relative ranking between the two architectures reverses, offering practical guidance for future near-memory system design.
title PIM or CXL-PIM? Understanding Architectural Trade-offs Through Large-Scale Benchmarking
topic Emerging Technologies
Performance
url https://arxiv.org/abs/2511.14400