Balanced Data Placement for GEMV Acceleration with Processing-In-Memory
Fuente:
arXiv
Saved in:
| Main Authors: | Ibrahim, Mohamed Assem, Islam, Mahzabeen, Aga, Shaizeen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
by: Pati, Suchita, et al.
Published: (2025)
by: Pati, Suchita, et al.
Published: (2025)
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
by: Singhania, Varsha, et al.
Published: (2024)
by: Singhania, Varsha, et al.
Published: (2024)
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024)
by: Agrawal, Anirudha, et al.
Published: (2024)
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
by: Pal, Shagnik, et al.
Published: (2025)
by: Pal, Shagnik, et al.
Published: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
Lit Silicon: A Case Where Thermal Imbalance Couples Concurrent Execution in Multiple GPUs
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Global Optimizations & Lightweight Dynamic Logic for Concurrency
by: Pati, Suchita, et al.
Published: (2024)
by: Pati, Suchita, et al.
Published: (2024)
Accelerating Triangle Counting with Real Processing-in-Memory Systems
by: Asquini, Lorenzo, et al.
Published: (2025)
by: Asquini, Lorenzo, et al.
Published: (2025)
A Modern Primer on Processing in Memory
by: Mutlu, Onur, et al.
Published: (2020)
by: Mutlu, Onur, et al.
Published: (2020)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
by: Yu, Yanpeng, et al.
Published: (2025)
by: Yu, Yanpeng, et al.
Published: (2025)
Memory-Centric Computing: Recent Advances in Processing-in-DRAM
by: Mutlu, Onur, et al.
Published: (2024)
by: Mutlu, Onur, et al.
Published: (2024)
CMDS: Cross-layer Dataflow Optimization for DNN Accelerators Exploiting Multi-bank Memories
by: Shi, Man, et al.
Published: (2024)
by: Shi, Man, et al.
Published: (2024)
New Tools, Programming Models, and System Support for Processing-in-Memory Architectures
by: Oliveira, Geraldo F.
Published: (2025)
by: Oliveira, Geraldo F.
Published: (2025)
NMP-PaK: Near-Memory Processing Acceleration of Scalable De Novo Genome Assembly
by: Kim, Heewoo, et al.
Published: (2025)
by: Kim, Heewoo, et al.
Published: (2025)
ALPHA-PIM: Analysis of Linear Algebraic Processing for High-Performance Graph Applications on a Real Processing-In-Memory System
by: Barkhordar, Marzieh, et al.
Published: (2026)
by: Barkhordar, Marzieh, et al.
Published: (2026)
PIMDAL: Mitigating the Memory Bottleneck in Data Analytics using a Real Processing-in-Memory System
by: Frouzakis, Manos, et al.
Published: (2025)
by: Frouzakis, Manos, et al.
Published: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026)
by: Liu, Lian, et al.
Published: (2026)
Accelerating Data Chunking in Deduplication Systems using Vector Instructions
by: Udayashankar, Sreeharsha, et al.
Published: (2025)
by: Udayashankar, Sreeharsha, et al.
Published: (2025)
Accelerating DNA Read Mapping with Digital Processing-in-Memory
by: Ben-Hur, Rotem, et al.
Published: (2024)
by: Ben-Hur, Rotem, et al.
Published: (2024)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
by: Chen, Yanru, et al.
Published: (2025)
by: Chen, Yanru, et al.
Published: (2025)
PIUMA: Programmable Integrated Unified Memory Architecture
by: Aananthakrishnan, Sriram, et al.
Published: (2020)
by: Aananthakrishnan, Sriram, et al.
Published: (2020)
EDEA: Efficient Dual-Engine Accelerator for Depthwise Separable Convolution with Direct Data Transfer
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Memory-Centric Computing: Solving Computing's Memory Problem
by: Mutlu, Onur, et al.
Published: (2025)
by: Mutlu, Onur, et al.
Published: (2025)
CIPHERMATCH: Accelerating Homomorphic Encryption-Based String Matching via Memory-Efficient Data Packing and In-Flash Processing
by: Kabra, Mayank, et al.
Published: (2025)
by: Kabra, Mayank, et al.
Published: (2025)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
by: Oliveira, Geraldo F., et al.
Published: (2024)
by: Oliveira, Geraldo F., et al.
Published: (2024)
XDMA: A Distributed, Extensible DMA Architecture for Layout-Flexible Data Movements in Heterogeneous Multi-Accelerator SoCs
by: Kong, Fanchen, et al.
Published: (2025)
by: Kong, Fanchen, et al.
Published: (2025)
Analyzing a Two-Tier Disaggregated Memory Protection Scheme Based on Memory Replication
by: Volos, Haris, et al.
Published: (2025)
by: Volos, Haris, et al.
Published: (2025)
Conduit: Programmer-Transparent Near-Data Processing Using Multiple Compute-Capable Resources in Solid State Drives
by: Nadig, Rakesh, et al.
Published: (2026)
by: Nadig, Rakesh, et al.
Published: (2026)
Sparse MTTKRP Acceleration for Tensor Decomposition on GPU
by: Wijeratne, Sasindu, et al.
Published: (2024)
by: Wijeratne, Sasindu, et al.
Published: (2024)
Leveraging SIMD for Accelerating Large-number Arithmetic
by: Das, Subhrajit, et al.
Published: (2026)
by: Das, Subhrajit, et al.
Published: (2026)
NasZip: Software and Hardware Co-Design to Accelerate Approximate Nearest Neighbor Search with DIMM-Based Near-Data Processing
by: Zou, Cheng, et al.
Published: (2026)
by: Zou, Cheng, et al.
Published: (2026)
Proteus: Enabling High-Performance Processing-Using-DRAM with Dynamic Bit-Precision, Adaptive Data Representation, and Flexible Arithmetic
by: Oliveira, Geraldo F., et al.
Published: (2025)
by: Oliveira, Geraldo F., et al.
Published: (2025)
Efficient Architecture for RISC-V Vector Memory Access
by: Guan, Hongyi, et al.
Published: (2025)
by: Guan, Hongyi, et al.
Published: (2025)
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
by: Zhang, Qijun, et al.
Published: (2026)
by: Zhang, Qijun, et al.
Published: (2026)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Handling of Memory Page Faults during Virtual-Address RDMA
by: Psistakis, Antonis
Published: (2025)
by: Psistakis, Antonis
Published: (2025)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
A Heterogeneous Chiplet Architecture for Accelerating End-to-End Transformer Models
by: Sharma, Harsh, et al.
Published: (2023)
by: Sharma, Harsh, et al.
Published: (2023)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
by: Elwasif, Wael, et al.
Published: (2022)
by: Elwasif, Wael, et al.
Published: (2022)
Similar Items
-
DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication
by: Pati, Suchita, et al.
Published: (2025) -
FinGraV: Methodology for Fine-Grain GPU Power Visibility and Insights
by: Singhania, Varsha, et al.
Published: (2024) -
Optimizing ML Concurrent Computation and Communication with GPU DMA Engines
by: Agrawal, Anirudha, et al.
Published: (2024) -
Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap
by: Pal, Shagnik, et al.
Published: (2025) -
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)