The Hitchhiker's Guide to Programming and Optimizing Cache Coherent Heterogeneous Systems: CXL, NVLink-C2C, and AMD Infinity Fabric
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zixuan, Mahar, Suyash, Li, Luyi, Park, Jangseon, Kim, Jinpyo, Michailidis, Theodore, Pan, Yue, Shen, Mingyao, Rosing, Tajana, Tullsen, Dean, Swanson, Steven, Zhao, Jishen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Formalising CXL Cache Coherence
von: Tan, Chengsong, et al.
Veröffentlicht: (2024)
von: Tan, Chengsong, et al.
Veröffentlicht: (2024)
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
von: Wang, Xuan, et al.
Veröffentlicht: (2024)
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026)
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)
CXL Shared Memory Programming: Barely Distributed and Almost Persistent
von: Xu, Yi, et al.
Veröffentlicht: (2024)
von: Xu, Yi, et al.
Veröffentlicht: (2024)
FaTRQ: Tiered Residual Quantization for LLM Vector Search in Far-Memory-Aware ANNS Systems
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026)
Efficient Open Modification Spectral Library Searching in High-Dimensional Space with Multi-Level-Cell Memory
von: Fan, Keming, et al.
Veröffentlicht: (2024)
von: Fan, Keming, et al.
Veröffentlicht: (2024)
Telepathic Datacenters: Fast RPCs using Shared CXL Memory
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
von: Mahar, Suyash, et al.
Veröffentlicht: (2024)
HAVEN: High-Bandwidth Flash Augmented Vector Engine for Large-Scale Approximate Nearest-Neighbor Search Acceleration
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
von: Hsu, Po-Kai, et al.
Veröffentlicht: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
von: Gouk, Donghyun, et al.
Veröffentlicht: (2025)
Hybrid SLC-MLC RRAM Mixed-Signal Processing-in-Memory Architecture for Transformer Acceleration via Gradient Redistribution
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
von: Song, Chang Eun, et al.
Veröffentlicht: (2025)
PREFENDER: A Prefetching Defender against Cache Side Channel Attacks as A Pretender
von: Li, Luyi, et al.
Veröffentlicht: (2023)
von: Li, Luyi, et al.
Veröffentlicht: (2023)
OneAdapt: Adaptive Compilation for Resource-Constrained Photonic One-Way Quantum Computing
von: Zhang, Hezi, et al.
Veröffentlicht: (2025)
von: Zhang, Hezi, et al.
Veröffentlicht: (2025)
Learning Cache Coherence Traffic for NoC Routing Design
von: Xiong, Guochu, et al.
Veröffentlicht: (2025)
von: Xiong, Guochu, et al.
Veröffentlicht: (2025)
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
von: Wang, Yanjing, et al.
Veröffentlicht: (2025)
von: Wang, Yanjing, et al.
Veröffentlicht: (2025)
CHIME: Chiplet-based Heterogeneous Near-Memory Acceleration for Edge Multimodal LLM Inference
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
von: Chen, Yanru, et al.
Veröffentlicht: (2025)
The Case for Persistent CXL switches
von: Hadi, Khan Shaikhul, et al.
Veröffentlicht: (2025)
von: Hadi, Khan Shaikhul, et al.
Veröffentlicht: (2025)
GEN-Graph: Heterogeneous PIM Accelerator for General Computational Patterns in Graph-based Dynamic Programming
von: Chen, Yanru, et al.
Veröffentlicht: (2026)
von: Chen, Yanru, et al.
Veröffentlicht: (2026)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
von: Xu, Weihong, et al.
Veröffentlicht: (2025)
Proxima: Near-storage Acceleration for Graph-based Approximate Nearest Neighbor Search in 3D NAND
von: Xu, Weihong, et al.
Veröffentlicht: (2023)
von: Xu, Weihong, et al.
Veröffentlicht: (2023)
ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
von: Chen, Hanqiu, et al.
Veröffentlicht: (2024)
von: Chen, Hanqiu, et al.
Veröffentlicht: (2024)
CXL-DMSim: A Full-System CXL Disaggregated Memory Simulator With Comprehensive Silicon Validation
von: Wang, Yanjing, et al.
Veröffentlicht: (2024)
von: Wang, Yanjing, et al.
Veröffentlicht: (2024)
RapidOMS: FPGA-based Open Modification Spectral Library Searching with HD Computing
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
von: Pinge, Sumukh, et al.
Veröffentlicht: (2024)
FedUHD: Unsupervised Federated Learning using Hyperdimensional Computing
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
von: Lee, You Hak, et al.
Veröffentlicht: (2025)
The Open-Source BlackParrot-BedRock Cache Coherence System
von: Wyse, Mark Unruh
Veröffentlicht: (2025)
von: Wyse, Mark Unruh
Veröffentlicht: (2025)
HDDB: Efficient In-Storage SQL Database Search Using Hyperdimensional Computing on Ferroelectric NAND Flash
von: Zhao, Quanling, et al.
Veröffentlicht: (2025)
von: Zhao, Quanling, et al.
Veröffentlicht: (2025)
Octopus: Enhancing CXL Memory Pods via Sparse Topology
von: Zhong, Yuhong, et al.
Veröffentlicht: (2025)
von: Zhong, Yuhong, et al.
Veröffentlicht: (2025)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
von: Wang, Jiapin, et al.
Veröffentlicht: (2024)
von: Wang, Jiapin, et al.
Veröffentlicht: (2024)
A Novel Extensible Simulation Framework for CXL-Enabled Systems
von: An, Yuda, et al.
Veröffentlicht: (2024)
von: An, Yuda, et al.
Veröffentlicht: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
von: Wang, Zhao, et al.
Veröffentlicht: (2025)
von: Wang, Zhao, et al.
Veröffentlicht: (2025)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
von: Kim, Dowon, et al.
Veröffentlicht: (2025)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
von: Oh, Dongsuk, et al.
Veröffentlicht: (2025)
von: Oh, Dongsuk, et al.
Veröffentlicht: (2025)
FlexLink: Boosting your NVLink Bandwidth by 27% without accuracy concern
von: Shen, Ao, et al.
Veröffentlicht: (2025)
von: Shen, Ao, et al.
Veröffentlicht: (2025)
Switchable Single/Dual Edge Registers for Pipeline Architecture
von: Singh, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Singh, Suyash Vardhan, et al.
Veröffentlicht: (2024)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
von: Mao, Shunyu, et al.
Veröffentlicht: (2024)
3D MPSoC with On-Chip Cache Support -- Design and Exploitation
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
von: Cataldo, Rodrigo, et al.
Veröffentlicht: (2025)
Rhea: a Framework for Fast Design and Validation of RTL Cache-Coherent Memory Subsystems
von: Zoni, Davide, et al.
Veröffentlicht: (2025)
von: Zoni, Davide, et al.
Veröffentlicht: (2025)
AMD Versal Implementations of FAM and SSCA Estimators
von: Li, Carol Jingyi, et al.
Veröffentlicht: (2025)
von: Li, Carol Jingyi, et al.
Veröffentlicht: (2025)
Performance Characterizations and Usage Guidelines of Samsung CXL Memory Module Hybrid Prototype
von: Zeng, Jianping, et al.
Veröffentlicht: (2025)
von: Zeng, Jianping, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Formalising CXL Cache Coherence
von: Tan, Chengsong, et al.
Veröffentlicht: (2024) -
Fast-OverlaPIM: A Fast Overlap-driven Mapping Framework for Processing In-Memory Neural Network Acceleration
von: Wang, Xuan, et al.
Veröffentlicht: (2024) -
SpANNS: Optimizing Approximate Nearest Neighbor Search for Sparse Vectors Using Near Memory Processing
von: Zhang, Tianqi, et al.
Veröffentlicht: (2026) -
GenDRAM:Hardware-Software Co-Design of General Platform in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2026) -
PIM-FW: Hardware-Software Co-Design of All-pairs Shortest Paths in DRAM
von: Lu, Tsung-Han, et al.
Veröffentlicht: (2025)