DCC: Data-Centric Compilation of Machine Learning Kernels for Processing-In-Memory Architectures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Peiming, Durvasula, Sankeerth, Fernandez, Ivan, Sadrosadati, Mohammad, Mutlu, Onur, Pekhimenko, Gennady, Giannoula, Christina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
von: Giannoula, Christina, et al.
Veröffentlicht: (2024)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025)
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
von: Mutlu, Onur, et al.
Veröffentlicht: (2020)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)
Compiler Support for Speculation in Decoupled Access/Execute Architectures
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
von: Szafarczyk, Robert, et al.
Veröffentlicht: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
von: Mutlu, Onur, et al.
Veröffentlicht: (2024)
Spira: Exploiting Voxel Data Structural Properties for Efficient Sparse Convolution in Point Cloud Networks
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
von: Adamopoulos, Dionysios, et al.
Veröffentlicht: (2025)
Memory-Centric Computing: Solving Computing's Memory Problem
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
von: Mutlu, Onur, et al.
Veröffentlicht: (2025)
RevaMp3D: Architecting the Processor Core and Cache Hierarchy for Systems with Monolithically-Integrated Logic and Memory
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2022)
Taking Cryptography Out of the Data Path via Near-Memory Processing in DRAM
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
von: Barcarolo, Nicola, et al.
Veröffentlicht: (2026)
PAPI: Exploiting Dynamic Parallelism in Large Language Model Decoding with a Processing-In-Memory-Enabled Computing System
von: He, Yintao, et al.
Veröffentlicht: (2025)
von: He, Yintao, et al.
Veröffentlicht: (2025)
Vector-Centric Machine Learning Systems: A Cross-Stack Approach
von: Jiang, Wenqi
Veröffentlicht: (2025)
von: Jiang, Wenqi
Veröffentlicht: (2025)
Dissecting the NVIDIA Hopper Architecture through Microbenchmarking and Multiple Level Analysis
von: Luo, Weile, et al.
Veröffentlicht: (2025)
von: Luo, Weile, et al.
Veröffentlicht: (2025)
Experimental Assessment of Containers Running on Top of Virtual Machines
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
von: Aqasizade, Hossein, et al.
Veröffentlicht: (2024)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
von: Qararyah, Fareed, et al.
Veröffentlicht: (2024)
Evaluating the Potential of In-Memory Processing to Accelerate Homomorphic Encryption
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2024)
CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash Processing
von: Kabra, Mayank, et al.
Veröffentlicht: (2025)
von: Kabra, Mayank, et al.
Veröffentlicht: (2025)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
von: Nadig, Rakesh, et al.
Veröffentlicht: (2026)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
von: Oliveira, Geraldo F., et al.
Veröffentlicht: (2025)
Performance Implications of Multi-Chiplet Neural Processing Units on Autonomous Driving Perception
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
von: Odema, Mohanad, et al.
Veröffentlicht: (2024)
GigaAPI for GPU Parallelization
von: Suvarna, M., et al.
Veröffentlicht: (2025)
von: Suvarna, M., et al.
Veröffentlicht: (2025)
Performance Analysis of HPC applications on the Aurora Supercomputer: Exploring the Impact of HBM-Enabled Intel Xeon Max CPUs
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
von: Ibeid, Huda, et al.
Veröffentlicht: (2025)
UPMEM Unleashed: Software Secrets for Speed
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
von: Chmielewski, Krystian, et al.
Veröffentlicht: (2025)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
von: Chakraborty, Abhinaba, et al.
Veröffentlicht: (2025)
Exploring and Evaluating Real-world CXL: Use Cases and System Adoption
von: Wang, Xi, et al.
Veröffentlicht: (2024)
von: Wang, Xi, et al.
Veröffentlicht: (2024)
Parallelizing a modern GPU simulator
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
von: Huerta, Rodrigo, et al.
Veröffentlicht: (2025)
Can Asymmetric Tile Buffering Be Beneficial?
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
von: Wang, Chengyue, et al.
Veröffentlicht: (2025)
Exploiting long vectors with a CFD code: a co-design show case
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
von: Blancafort, Marc, et al.
Veröffentlicht: (2024)
Simopt -- Simulation pass for Speculative Optimisation of FPGA-CAD flow
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
von: Wadhwa, Eashan, et al.
Veröffentlicht: (2024)
SAGe: A Lightweight Algorithm-Architecture Co-Design for Mitigating the Data Preparation Bottleneck in Large-Scale Genome Sequence Analysis
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2025)
GPT-OSS-20B: A Comprehensive Deployment-Centric Analysis of OpenAI's Open-Weight Mixture of Experts Model
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
von: Kumar, Deepak, et al.
Veröffentlicht: (2025)
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System
von: Rhyner, Steve, et al.
Veröffentlicht: (2024)
von: Rhyner, Steve, et al.
Veröffentlicht: (2024)
Optimizing Attention on GPUs by Exploiting GPU Architectural NUMA Effects
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
von: Choudhary, Mansi, et al.
Veröffentlicht: (2025)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
von: Vellaisamy, Prabhu, et al.
Veröffentlicht: (2025)
Massimult: A Novel Parallel CPU Architecture Based on Combinator Reduction
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
von: Nicklisch-Franken, Jurgen, et al.
Veröffentlicht: (2024)
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2024)
von: Ghiasi, Nika Mansouri, et al.
Veröffentlicht: (2024)
Minuet: Accelerating 3D Sparse Convolutions on GPUs
von: Yang, Jiacheng, et al.
Veröffentlicht: (2023)
von: Yang, Jiacheng, et al.
Veröffentlicht: (2023)
ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
von: Fan, Ruibo, et al.
Veröffentlicht: (2026)
Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry
von: Chen, Ziji, et al.
Veröffentlicht: (2025)
von: Chen, Ziji, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PyGim: An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures
von: Giannoula, Christina, et al.
Veröffentlicht: (2024) -
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
von: Barkhordar, Marzieh, et al.
Veröffentlicht: (2026) -
Accelerating Triangle Counting with Real Processing-in-Memory Systems
von: Asquini, Lorenzo, et al.
Veröffentlicht: (2025) -
A Modern Primer on Processing in Memory
von: Mutlu, Onur, et al.
Veröffentlicht: (2020) -
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
von: Frouzakis, Manos, et al.
Veröffentlicht: (2025)