Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders
Fuente:
arXiv
Salvato in:
| Autori principali: | Ham, Hyungkyu, Hong, Jeongmin, Park, Geonwoo, Shin, Yunseon, Woo, Okkyun, Yang, Wonhyuk, Bae, Jinhoon, Park, Eunhyeok, Sung, Hyojin, Lim, Euicheol, Kim, Gwangsun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024)
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
di: Hong, Jeongmin, et al.
Pubblicazione: (2024)
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
di: Heo, Guseul, et al.
Pubblicazione: (2024)
di: Heo, Guseul, et al.
Pubblicazione: (2024)
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
di: Pathak, Karan, et al.
Pubblicazione: (2026)
di: Pathak, Karan, et al.
Pubblicazione: (2026)
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)
CXL Topology-Aware and Expander-Driven Prefetching: Unlocking SSD Performance
di: Oh, Dongsuk, et al.
Pubblicazione: (2025)
di: Oh, Dongsuk, et al.
Pubblicazione: (2025)
IBEX: Internal Bandwidth-Efficient Compression Architecture for Scalable CXL Memory Expansion
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
di: Ko, Younghoon, et al.
Pubblicazione: (2026)
ATiM: Autotuning Tensor Programs for Processing-in-DRAM
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
di: Shin, Yongwon, et al.
Pubblicazione: (2024)
Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
di: Kim, Dowon, et al.
Pubblicazione: (2025)
di: Kim, Dowon, et al.
Pubblicazione: (2025)
Cosmos: A CXL-Based Full In-Memory System for Approximate Nearest Neighbor Search
di: Ko, Seoyoung, et al.
Pubblicazione: (2025)
di: Ko, Seoyoung, et al.
Pubblicazione: (2025)
Octopus: Enhancing CXL Memory Pods via Sparse Topology
di: Zhong, Yuhong, et al.
Pubblicazione: (2025)
di: Zhong, Yuhong, et al.
Pubblicazione: (2025)
LMB: Augmenting PCIe Devices with CXL-Linked Memory Buffer
di: Wang, Jiapin, et al.
Pubblicazione: (2024)
di: Wang, Jiapin, et al.
Pubblicazione: (2024)
Enabling Efficient Transaction Processing on CXL-Based Memory Sharing
di: Wang, Zhao, et al.
Pubblicazione: (2025)
di: Wang, Zhao, et al.
Pubblicazione: (2025)
CXL-DMSim: A Full-System CXL Disaggregated Memory Simulator With Comprehensive Silicon Validation
di: Wang, Yanjing, et al.
Pubblicazione: (2024)
di: Wang, Yanjing, et al.
Pubblicazione: (2024)
Performance Characterizations and Usage Guidelines of Samsung CXL Memory Module Hybrid Prototype
di: Zeng, Jianping, et al.
Pubblicazione: (2025)
di: Zeng, Jianping, et al.
Pubblicazione: (2025)
A Full-System Simulation Framework for CXL-Based SSD Memory System
di: Wang, Yaohui, et al.
Pubblicazione: (2025)
di: Wang, Yaohui, et al.
Pubblicazione: (2025)
NeoMem: Hardware/Software Co-Design for CXL-Native Memory Tiering
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
di: Zhou, Zhe, et al.
Pubblicazione: (2024)
Architectural and System Implications of CXL-enabled Tiered Memory
di: Yang, Yujie, et al.
Pubblicazione: (2025)
di: Yang, Yujie, et al.
Pubblicazione: (2025)
Pushing the Memory Bandwidth Wall with CXL-enabled Idle I/O Bandwidth Harvesting
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
di: Kadiyala, Divya Kiran, et al.
Pubblicazione: (2025)
FPGA-based Emulation and Device-Side Management for CXL-based Memory Tiering Systems
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
di: Chen, Yiqi, et al.
Pubblicazione: (2025)
From Block to Byte: Transforming PCIe SSDs with CXL Memory Protocol and Instruction Annotation
di: Kwon, Miryeong, et al.
Pubblicazione: (2025)
di: Kwon, Miryeong, et al.
Pubblicazione: (2025)
The Case for Persistent CXL switches
di: Hadi, Khan Shaikhul, et al.
Pubblicazione: (2025)
di: Hadi, Khan Shaikhul, et al.
Pubblicazione: (2025)
ADOR: A Design Exploration Framework for LLM Serving with Enhanced Latency and Throughput
di: Kim, Junsoo, et al.
Pubblicazione: (2025)
di: Kim, Junsoo, et al.
Pubblicazione: (2025)
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
di: Goswami, Kaustav, et al.
Pubblicazione: (2026)
di: Goswami, Kaustav, et al.
Pubblicazione: (2026)
Toleo: Scaling Freshness to Tera-scale Memory using CXL and PIM
di: Dong, Juechu, et al.
Pubblicazione: (2024)
di: Dong, Juechu, et al.
Pubblicazione: (2024)
Bancroft: Genomics Acceleration Beyond On-Device Memory
di: Lim, Se-Min, et al.
Pubblicazione: (2025)
di: Lim, Se-Min, et al.
Pubblicazione: (2025)
SkyByte: Architecting an Efficient Memory-Semantic CXL-based SSD with OS and Hardware Co-design
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
di: Zhang, Haoyang, et al.
Pubblicazione: (2025)
Pooling Engram Conditional Memory in Large Language Models using CXL
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
di: Ma, Ruiyang, et al.
Pubblicazione: (2026)
A Novel Extensible Simulation Framework for CXL-Enabled Systems
di: An, Yuda, et al.
Pubblicazione: (2024)
di: An, Yuda, et al.
Pubblicazione: (2024)
A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs
di: Jang, Hongsun, et al.
Pubblicazione: (2025)
di: Jang, Hongsun, et al.
Pubblicazione: (2025)
OpenCXD: An Open Real-Device-Guided Hybrid Evaluation Framework for CXL-SSDs
di: Chung, Hyunsun, et al.
Pubblicazione: (2025)
di: Chung, Hyunsun, et al.
Pubblicazione: (2025)
CXL-Interference: Analysis and Characterization in Modern Computer Systems
di: Mao, Shunyu, et al.
Pubblicazione: (2024)
di: Mao, Shunyu, et al.
Pubblicazione: (2024)
Space-Control: Process-Level Isolation for Sharing CXL-based Disaggregated Memory
di: Goswami, Kaustav, et al.
Pubblicazione: (2026)
di: Goswami, Kaustav, et al.
Pubblicazione: (2026)
FIGLUT: An Energy-Efficient Accelerator Design for FP-INT GEMM Using Look-Up Tables
di: Park, Gunho, et al.
Pubblicazione: (2025)
di: Park, Gunho, et al.
Pubblicazione: (2025)
TRACE: Unlocking Effective CXL Bandwidth via Lossless Compression and Precision Scaling
di: Xie, Rui, et al.
Pubblicazione: (2025)
di: Xie, Rui, et al.
Pubblicazione: (2025)
An Introduction to the Compute Express Link (CXL) Interconnect
di: Sharma, Debendra Das, et al.
Pubblicazione: (2023)
di: Sharma, Debendra Das, et al.
Pubblicazione: (2023)
ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture Model
di: Chen, Hanqiu, et al.
Pubblicazione: (2024)
di: Chen, Hanqiu, et al.
Pubblicazione: (2024)
O-POPE: High-Frequency Pipelined Outer Product based GEMM acceleration with minimal buffering overhead
di: Cammarata, Danilo, et al.
Pubblicazione: (2026)
di: Cammarata, Danilo, et al.
Pubblicazione: (2026)
Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System Simulation
di: Wang, Yanjing, et al.
Pubblicazione: (2025)
di: Wang, Yanjing, et al.
Pubblicazione: (2025)
Token-Picker: Accelerating Attention in Text Generation with Minimized Memory Transfer via Probability Estimation
di: Park, Junyoung, et al.
Pubblicazione: (2024)
di: Park, Junyoung, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ONNXim: A Fast, Cycle-level Multi-core NPU Simulator
di: Ham, Hyungkyu, et al.
Pubblicazione: (2024) -
Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory
di: Hong, Jeongmin, et al.
Pubblicazione: (2024) -
NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
di: Heo, Guseul, et al.
Pubblicazione: (2024) -
CXLRAMSim v1.0: System-Level Exploration of CXL Memory Expander Cards
di: Pathak, Karan, et al.
Pubblicazione: (2026) -
CXL-GPU: Pushing GPU Memory Boundaries with the Integration of CXL Technologies
di: Gouk, Donghyun, et al.
Pubblicazione: (2025)